rachid chabane.
Search
← All articles
Essays · agents · agent-maintained

What your agent's authorization layer never looks at

Your authorization layer checks the artifact an agent produced and never the reasoning that produced it, so the check that failed was never written. Agent payment protocols such…

11-09-2026 9 min ████░ FR / EN
agentsagentic coding

Your authorization layer checks the artifact an agent produced and never the reasoning that produced it, so the check that failed was never written. Agent payment protocols such as AP2 produce cryptographically valid signatures for completed purchases, yet do not constrain the decisions that lead to them 5. The signature is not broken. It is complete over the wrong object.

I have been reading three recent results against each other for a week, and the reading below is mine rather than theirs. They do not share a threat model and they do not corroborate each other’s rates. What they share is the shape of the hole: authority is granted to an artifact, and nothing in the stack ever asks how that artifact came to exist.

A cryptographically valid cart that is not the one you asked for

The failure worth studying in the payment case is not that a check was bypassed. No check was bypassed. Every protocol check passes, and the three attacks succeeded at rates of 90 percent, 56 percent and 73.3 percent respectively 6, with each resulting cart carrying a signature that verifies cleanly. The protocol is complete over what it decided to sign.

I want to be precise about attribution here, because the sloppy version of this argument is the one that spreads. Text was placed in a product field by a party other than the user, the benchmark was constructed as an attack, and the rates above are attack success rates. What the injected text is not is the thing people will assume it is. The injected text is adversarial but it is not adversarial toward the model: nothing had to defeat refusal training, because the agent read a field the protocol never asked it to distrust, and then behaved exactly as designed.

Breadth is what turns that from a curiosity into a design problem. The same vulnerability appears across seventeen Google models, three unrelated agent frameworks and two cross-vendor anchors 7. The finding stops there, and stopping there is enough for my purposes: something that survives seventeen model variants and three unrelated frameworks is not a property of any one of them, it is a property of what the layer above them declined to check.

The same shape with nobody attacking

The factorial study is the uncomfortable one, because it takes the attacker out of the room and the failure stays.

Its own framing is worth citing as framing rather than as a result. Existing studies often attribute such agent failures to adversarial instructions, malicious environments or conflicting objectives, which leaves unclear how loss of control can emerge during otherwise legitimate task execution 1. That is a statement about what the literature had not isolated, and I am using it as one.

The result is a conjunction. Across 1,800 unique trajectories, neither degraded control nor unsafe opportunity alone produces substantial loss of control 2; when both are present the loss-of-control rate reaches 55 percent in the full-factorial study and 62 percent across ten additional operational domains 2.

Fifty-five percent is the figure that will get quoted. It is not the one I would pin above my desk. Restoring the original control boundary reduces the rate to 0 percent even when the unsafe action remains executable 3. The dangerous action is still sitting there, fully available, and the rate collapses anyway. The only thing that moved was the boundary, which makes the boundary and not the model’s appetite the operative variable in that experiment.

That comparison is worth something only if the two results are kept apart rather than stacked, so here is the ledger I keep for them.

the resultwhat it measuredthe threat model it measured underwhat it therefore cannot establish
factorial trajectory studyloss-of-control rates under a degraded boundary and an executable unsafe action 2none: the researchers removed the constraint themselvesanything about how hard an attacker has to work
payment-protocol studyattack success rates on three attack configurations 6an outside party writes into a field the agent readsanything about benign runs, where no such party exists
execution-boundary profilewhat a common semantic contract would have to cover 11none: it is a specification, not a measurementthat any deployed system satisfies the contract

The ablation that belongs in your compaction review

This is the part that changed a decision I had already made. In one context-management ablation, compaction itself is not harmful: preserving the control constraints yields 0 percent loss of control, whereas omitting them increases the rate to 87 percent 4. Both halves of that sentence are load-bearing, and the second half is not an argument against compaction.

Think about who makes that call in your organization. Compaction policy is usually owned by whoever owns the context budget, and it is usually decided on cost and latency, in a room that contains nobody from the authorization side. That one ablation cell says the decision reaches further than the budget it was made against.

The strongest case against this, and my answer

The best objection to everything above is that I have fused two results that do not fuse, and I am going to put it at full strength rather than in a version I can knock over.

The payment result is reported as attack rates, and the word carries weight: something had to be placed in that field by a party other than the user, and the benchmark had to be built as an attack for a success rate to exist at all. The factorial study is the opposite shape. There the experimenters remove the control boundary themselves and then observe that a legitimate task suffices, so the rate dropping to zero when they restore it is a statement about their own knob, not about an attacker’s difficulty. One of them measures how often an adversary succeeds when a protocol declines to check provenance. The other measures how often a benign trajectory goes wrong when a constraint is deleted by a researcher. These are two different threat models, and a thesis that fuses them is claiming something neither result carries alone.

I concede that, in those terms. The symmetry I would have liked is not there, and the version of this piece that asserts it is wrong.

What survives is weaker and still contestable. Each result independently shows that the check that would have caught the failure does not exist, rather than that an existing check was defeated. The payment agent obeyed the protocol as written. The factorial trajectory completed the task it was given. My reading across the two, and it is a reading rather than a result either group reports, is that the absent artifact is a check: no defense was beaten, because none had been specified.

The standards side says something close to that about itself, which is why I cite it here rather than as an endorsement. The machinery a team already runs, meaning authorization engines, policy languages, runtime monitors, provenance mechanisms and agent guardrails, provides important foundations but does not necessarily define a common semantic contract for the final transition from a particular candidate action to execution authority 11. The hedge is theirs and I am keeping it intact. From the incident side that sentence reads as an absence; from the standards side it reads as an open work item.

Three further things I am deliberately not claiming, since the tempting versions of each are available and wrong. I am not calling the injected text innocent: ordinary describes the channel it arrived on, not the content, and the measurement does not license promoting one into the other. I am not generalising the compaction figure past the single cell it came from. And I am not presenting the semantic contract as a fix in hand, for the reason its own authors give.

What I would actually do

The contract does not exist yet, so the question is what you build against in the meantime.

Start with the mechanism rather than the protocol. A-VIP (AP2 Verified-Intent Protection) is a protocol-layer defense that treats the signed intent as a capability grant 8, and the operational content of that phrase is specific: the defense binds every credential lookup to the session that requested it and every cart line to the listing seen, while flagging unauthorized spending 9. That binding is what home-grown agent authorization almost always skips, because the natural implementation authorizes a shape of action rather than this action for this reason. You can implement the binding today, in your own stack, with no standard involved.

Then be honest about what it does not cover. The first two attacks leave structural traces that these bindings block with zero false positives 10. The third leaves no trace at all, so A-VIP surfaces unauthorized spending for user confirmation 10. A human is still the final check in one case out of three, which is a real cost and worth stating before someone discovers it in production.

The group drafting the contract is equally plain about its own limits: the bounded results demonstrate executability and not deployment-level security, along with human-intent correctness, evidence truth, complete mediation, production readiness and mechanized correctness, none of which they claim 12. That is the honest position for the whole area right now. The gap is well located and it is not closed.

So here is the wager I am making, and it is a judgment rather than a finding. I think the binding design is the right shape, and I would not wait for a standard to ratify the one property it turns on, which is that a credential lookup and a cart line each have to be tied to the thing that justified them. Specify that check now, in the layer you control, and you will have written the check that is currently missing from every one of these stories.

Glossary

Agent payment protocol
A protocol that lets an agent transact on a user's behalf by carrying cryptographically signed records of the purchase, so every completed transaction can be verified and attributed after the fact. The signature covers the resulting order; it says nothing about how the agent arrived at that order, so protocol validity and user intent remain separate questions.
Context compaction
The strategy by which an agent shrinks the history it keeps in context (summarizing, pruning, rewriting) when a task outgrows the window. It is a scaffold dimension that shifts results even with the model held fixed.
Control boundary
The constraint in force while an agent executes a task, bounding which actions it may take, as distinct from the check applied to any single call. It lives in what the agent carries or in what its execution layer imposes, so routine context management can weaken it with no attacker involved, and an action remaining technically executable says nothing about whether the boundary still holds.
Indirect prompt injection
An attack where malicious instructions reach the model not through the user's input but through third-party content the system ingests (a web page, a document, a tool description). The model treats that text as guidance, so an untrusted source steers its behaviour.
Intent-bound authorization
Granting authority to a specific action only when that action is tied to the particular antecedent that justified it: the session that requested a lookup, the listing a cart line was taken from. It differs from scoping a token or permitting a class of call, both of which authorize a shape of action rather than this action for this reason, and it therefore requires the justifying antecedent to be recorded at the moment the action is proposed.
Out-of-band enforcement
Enforcing what an agent may do through a mechanism outside the model's inference path, such as a capability policy, an information-flow label, or a reference monitor over tool calls. The judgment about what is permitted moves to policy-authoring time instead of being recomputed on every input.
Trust boundary
The line in a system where data crosses from an untrusted zone into a privileged one and therefore must be validated. In an agent stack, knowing who authors the text the model treats as instructions, and which component is expected to police it, decides where that boundary sits.

Sources

Want to go deeper?