<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Security on Unladen swallow - Olivier Wulveryck</title>
    <link>https://blog.owulveryck.info/tags/security.html</link>
    <description>Recent content in Security on Unladen swallow - Olivier Wulveryck</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <copyright>Olivier Wulveryck</copyright>
    <lastBuildDate>Thu, 08 Oct 2026 10:00:00 +0200</lastBuildDate>
    
        <atom:link href="https://blog.owulveryck.info/tags/security/index.xml" rel="self" type="application/rss+xml" />
    
    
    <item>
      <title>The Agentic Threat</title>
      <link>https://blog.owulveryck.info/2026/10/08/the-agentic-threat.html</link>
      <pubDate>Thu, 08 Oct 2026 10:00:00 +0200</pubDate>
      
      <guid>https://blog.owulveryck.info/2026/10/08/the-agentic-threat.html</guid>
      
        <description>&lt;h2 id=&#34;autonomy-is-convenient&#34;&gt;Autonomy is convenient&lt;/h2&gt;
&lt;p&gt;I am more than convinced that &lt;strong&gt;agentic engineering&lt;/strong&gt; is the discipline that will unlock the full value of AI in the enterprise.
I am lucky enough to work on the evolution of the processes used to build digital assets, software and others.
In that context, one key is to provide agentic systems that act with &lt;strong&gt;as much autonomy as possible&lt;/strong&gt; to carry out non-differentiating tasks robustly and quickly (and ideally at a controlled cost).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Development harnesses&lt;/strong&gt; have opened the way to agent autonomy. It is now possible to have a conversation to create digital assets with Claude Code, Copilot or others.
You switch on an auto-mode and, after some framing, off it goes.
The harness comes with a set of tools that lets it interact with the ecosystem. In a company, interacting with the ecosystem means retrieving information, acting on tools or on processes.&lt;/p&gt;
&lt;p&gt;The more tools we give, the more autonomously agents can work. And as I was saying to a colleague this morning: the more autonomy we give, the more convenient it is…. And &lt;strong&gt;&amp;ldquo;it&amp;rsquo;s convenient&amp;rdquo; is one of the most dangerous phrases&lt;/strong&gt; in the digital ecosystem…. Because systems that offer convenience usually do so in exchange for something else of value to themselves.
In the case of agentic systems, we &lt;strong&gt;trade task delegation for control&lt;/strong&gt;: we hand autonomy over to the agent, and we widen the &lt;strong&gt;attack surface&lt;/strong&gt; accordingly.&lt;/p&gt;
&lt;p&gt;And that autonomy can be dangerous.&lt;/p&gt;
&lt;h2 id=&#34;two-risks-two-threats&#34;&gt;Two risks, two threats&lt;/h2&gt;
&lt;p&gt;We know that integrating software carries &lt;strong&gt;two risks&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data corruption&lt;/strong&gt; (the agent breaks everything), a database or a codebase for example.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data disclosure&lt;/strong&gt;: the agent retrieves data and may be programmed to transfer it to third parties, or simply hand it to its malicious pilot.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And these risks are carried by &lt;strong&gt;two threats&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The &lt;strong&gt;human threat&lt;/strong&gt;: the agent is instructed to do things it is not allowed to do, either to sabotage systems (corruption) or for espionage (disclosure). And the pilot is not necessarily a stranger on the Internet: it can be a &lt;strong&gt;malicious or manipulated employee&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;threat of the agent itself&lt;/strong&gt;: the agent is intrinsically stupid and could corrupt or disclose data &amp;ldquo;by accident&amp;rdquo;. I will set aside for now the threat of the &lt;strong&gt;&amp;ldquo;double&amp;rdquo; agent&lt;/strong&gt;, which would have a secret, learned intention leading it to deliberately extract or corrupt data (to hide a backdoor, for example). I will just note that, seen from the outside, &lt;strong&gt;a stupid agent and a double agent do the same thing&lt;/strong&gt;: what we control for one partly protects us from the other.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Addressing this requires agentic engineering, but in terms of architecture, &lt;strong&gt;these two threats do not have the same answer&lt;/strong&gt;.&lt;/p&gt;
&lt;h2 id=&#34;the-human-threat-ubers-persona-guardrail&#34;&gt;The human threat: Uber&amp;rsquo;s Persona Guardrail&lt;/h2&gt;
&lt;p&gt;For the human threat, one answer has been provided by Uber in its paper &lt;em&gt;Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems&lt;/em&gt; (ref &lt;a href=&#34;https://arxiv.org/abs/2610.03434&#34;&gt;arxiv 2610.03434&lt;/a&gt;). The paper is long and I admit I had an LLM ingest it to grasp its essence. The idea is to put a first LLM between the pilot and the agent to check that the request falls within the agent&amp;rsquo;s &lt;strong&gt;declared scope&lt;/strong&gt; (a list of allowed intents), and thus get a green light to start execution. The same LLM checks the &lt;strong&gt;final answer&lt;/strong&gt; before it goes back to the pilot. However, &lt;strong&gt;what happens in between (tool calls, changes) is not controlled&lt;/strong&gt;, and the paper says so itself.&lt;/p&gt;
&lt;h2 id=&#34;protecting-the-agent-from-itself-ppg-and-jevgo&#34;&gt;Protecting the agent from itself: PPG and jevgo&lt;/h2&gt;
&lt;p&gt;On the other hand, I think we also need &lt;strong&gt;more deterministic validations&lt;/strong&gt; at the level of the &lt;strong&gt;agentic loop&lt;/strong&gt; to protect the agent from itself. This is what I am exploring with &lt;strong&gt;PPG&lt;/strong&gt; (&lt;a href=&#34;https://github.com/owulveryck/poc-agentic-platform&#34;&gt;poc-agentic-platform&lt;/a&gt;): a gateway that validates, with OPA/Rego rules, the agent&amp;rsquo;s plan, each of its changes and the final diff. &lt;strong&gt;No ticket, no change.&lt;/strong&gt;
I have also considered setting up a &lt;strong&gt;&amp;ldquo;system 1&amp;rdquo;&lt;/strong&gt; (in Kahneman&amp;rsquo;s sense: fast, statistical, intuitive) next to the deterministic loop, with &lt;strong&gt;classification models&lt;/strong&gt; like &lt;a href=&#34;https://en.wikipedia.org/wiki/Jev_%28AI_model%29&#34;&gt;Jev&lt;/a&gt; (and my toy implementation: &lt;a href=&#34;https://github.com/owulveryck/jevgo&#34;&gt;jevgo&lt;/a&gt;): small models that learn to imitate the rules. &lt;strong&gt;They decide nothing.&lt;/strong&gt; When their verdict diverges from the rules&amp;rsquo;, it is a sign that a &lt;strong&gt;rule may be badly written&lt;/strong&gt;, and we escalate when in doubt.&lt;/p&gt;
&lt;h2 id=&#34;compensating-and-amplifying-measures&#34;&gt;Compensating and amplifying measures&lt;/h2&gt;
&lt;p&gt;To read these measures, I make a distinction:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;compensating&lt;/strong&gt; measure compensates for a weakness: it blocks, and each block requires a human to step in;&lt;/li&gt;
&lt;li&gt;an &lt;strong&gt;amplifying&lt;/strong&gt; measure lets the agentic loop correct itself: the error goes back to the agent, which fixes it before a human has to step in. The human moves from &lt;strong&gt;&amp;ldquo;in-the-loop&amp;rdquo;&lt;/strong&gt; (validating every step) to &lt;strong&gt;&amp;ldquo;on-the-loop&amp;rdquo;&lt;/strong&gt; (supervising and handling exceptions).&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;three-infographics&#34;&gt;Three infographics&lt;/h2&gt;
&lt;p&gt;Here is a summary in three infographics.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Where each guardrail acts.&lt;/strong&gt; Uber filters what is asked and what is answered; PPG controls what the agent does in between.&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://blog.owulveryck.info/assets/menace-agentique/where-guardrails-act.en.svg&#34; alt=&#34;Where each guardrail acts in the life of a request&#34; fetchpriority=&#34;high&#34; decoding=&#34;async&#34; /&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Same principle, opposite mechanisms.&lt;/strong&gt; Both declare a scope and deny by default. But you cannot write an exact rule for natural language, whereas you can for a plan or a diff.&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://blog.owulveryck.info/assets/menace-agentique/same-principle-opposite-mechanisms.en.svg&#34; alt=&#34;Same principle, opposite mechanisms: the shape of the input decides&#34; loading=&#34;lazy&#34; decoding=&#34;async&#34; /&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. System 1 observes, it does not judge.&lt;/strong&gt; jevgo runs next to the rules. A disagreement goes up to a human, who fixes the rule once and for all.&lt;/p&gt;
&lt;p&gt;&lt;img src=&#34;https://blog.owulveryck.info/assets/menace-agentique/jevgo-observer.en.svg&#34; alt=&#34;Adding jevgo to PPG: an observer of the rules, not a judge&#34; loading=&#34;lazy&#34; decoding=&#34;async&#34; /&gt;&lt;/p&gt;
&lt;h2 id=&#34;in-conclusion&#34;&gt;In conclusion&lt;/h2&gt;
&lt;p&gt;Uber&amp;rsquo;s architectural measure is &lt;strong&gt;essentially compensating&lt;/strong&gt; (its improvement loop makes the policy progress, not the agent). And it &lt;strong&gt;does not compensate for the agent&amp;rsquo;s stupidity but for its obedience&lt;/strong&gt;: a smarter agent will obey a malicious request better, so the measure will not become useless as models improve. On the contrary, it may struggle to keep up. The &lt;strong&gt;classifier is deliberately small&lt;/strong&gt; to stay under 100 ms. The better agents understand innuendo, the more a subtle request can be understood by the agent without being caught by the classifier, and the more &lt;strong&gt;false negatives&lt;/strong&gt; the gateway will let through.
On the PPG side, there is obviously a compensating aspect, since I stated that the primary goal was to address the stupidity of models. But there is also an &lt;strong&gt;amplifying aspect&lt;/strong&gt; for the agentic loop, which lets it self-correct when the gateway returns an error, before a human notices. The human no longer validates every step; they stay &lt;strong&gt;on-the-loop&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;In any case, a real agentic architecture of the future must take these two threats into account, by &lt;strong&gt;combining measures that protect and measures that make autonomy safe&lt;/strong&gt;.&lt;/p&gt;
</description>
      
    </item>
    
  </channel>
</rss>
