Selected for GitHub's Secure Open Source Fund. See how it's shaping the future of AI agent security.

Learn more

Blogs

Nobody Can Prove Who Authorised the Agent. We Checked All Thirteen Protocols.

Nasiko

Cover art: a figure stands in the dark before a chain of four glowing gold panels linked by check marks, a padlock on the line beneath the first.

We scored thirteen inter-agent protocols on one ten-axis rubric, with a recorded reason for every cell. This is what that found. The full survey is attached.

Every protocol comparison we could find was a narrative. Protocol A emphasises discovery, protocol B has a richer task model, both are promising, more work is needed. Four broad surveys exist and they are all useful. None of them lets you check the comparison.

So we built a rubric instead. Ten axes, ordinal levels, and each level anchored to a property you can find in a specification rather than to a judgement about quality. Then we read thirteen specifications and scored eight of them against all ten axes. For each of those eighty cells we recorded the sentence in the spec that put it at that level.

The point of anchoring the levels is that the scores stop being opinions. If you think A2A deserves better than level 1 on authorisation, you are not disagreeing with our taste. You are claiming a specific clause exists, and we can both go look.

What the rubric found on its third axis

Axis 3 asks how authorisation propagates. Agent A asks agent B to do something on behalf of principal P. Can B prove that P authorised this?

The levels run 0 to 3. Level 1 is OAuth scopes or token forwarding: A hands B its own token. Level 2 is a signed capability with explicit scope, audience, and time bounds, binding two named parties. Level 3 is a chain, where each hop's authority is independently verifiable by an agent that did not witness the grant.

Level 3 is empty. Not sparse. Empty.

Heatmap: eight protocols — A2A, MCP, ANP, ACP, AGNTCY/SLIM, Coral, LOKA, ACNBP — scored on discovery, identity, auth/delegation, task model, async, multi-modal, security (of 12) and semantic interop. The auth/delegation column holds only 1s and 2s; level 3 has no occupants.
Eight protocols scored. The top level of the authorisation axis has no occupants.

Four of the eight sit at level 1 and four at level 2. Not one specifies a delegation chain a third party can check. Every one of them leaves the mechanism to "the surrounding identity architecture," which is a specification-layer way of saying somebody else's problem.

The commerce protocols do reach level 3, which is the detail that makes this worth stating carefully. Their intent, mandate, and receipt envelope gets there by construction, because payment authorisation was never going to work any other way. So the field has a working example of the thing the peer-messaging layer lacks. It just sits in a layer that only moves money.

We went looking for the counterexample, because a universal claim deserves an honest attempt to break it. The best candidate is ACNBP. It has the widest threat coverage in the set, ten of twelve classes addressed, against one of twelve for A2A and MCP. It has a Capability Binding ceremony, which is the most rigorous single-hop primitive we read anywhere. It still binds exactly two parties at a time. A lead author confirmed that this is the intended scope, which is the kind of answer that makes a finding stronger rather than weaker.

Rigour on one hop is not authority that travels. Those are different properties, and the rubric had to separate them before we could see that no protocol has the second one.

Why an empty cell is worth a paper

Because deployments do not wait for specifications. They need delegation to work today, so they forward tokens, and the security literature has already measured what that costs.

The AgentLeak benchmark is the number that reframed this for us. Inter-agent messages leak personally identifiable content at 68.8%. The output channel, the one everybody audits, leaks at 27.2%. Output-only auditing misses roughly 41.7% of violations.

So the gap is an omission at the specification layer with a measured cost at the deployment layer. That is more actionable than "authorisation needs more work." It is the one item we would put at the top of the next standardisation cycle's agenda.

The result we did not expect

Scatter plot titled “The two strongest specifications have no documented deployment”: ACNBP (24) and LOKA (23) sit at zero publicly documented deployments, while A2A (14) and MCP (10) have the most deployments, around eleven each.

We assumed rubric strength would track adoption, at least loosely. It does not. It runs the other way.

The two highest-scoring protocols, ACNBP at 24 points and LOKA at 23, have the smallest deployment footprints we could document. The two with the largest footprints, A2A at 14 and MCP at 10, have modest profiles. A2A and MCP each address exactly one of twelve threat classes normatively.

What predicts adoption is institutional position. Foundation stewardship, vendor backing, ecosystem support. Technical rigour at the specification level predicts very little.

We report this as a finding rather than a complaint, because it has a consequence for anyone doing this work. The highest-impact technical contribution is the one an institution is positioned to adopt. That is an argument for working inside the AAIF and W3C tracks rather than alongside them. We took our own advice.

MCP versus A2A is a category error

Table of three layers: Layer 1 Tool Binding (MCP, ACP), Layer 2 Peer Messaging (A2A, ANP, AGNTCY/SLIM, Coral, LOKA, ACNBP), Layer 3 Commerce Settlement (AP2, Visa TAP, Mastercard Agent Pay, ERC-8004, UCP). MCP binds an agent to its tools, A2A carries messages between peers; they compose rather than compete.

The most-debated rivalry in the field is not a rivalry. MCP is a tool-binding protocol, A2A is a peer-messaging protocol, and they occupy different layers of the taxonomy. In production they compose. They do not substitute.

The stack the rubric most clearly recommends is not deployed anywhere: ANP's identity layer underneath A2A's task layer. ANP tops the identity axis on DIDs and verifiable credentials. A2A tops the task axis. No surveyed protocol combines a deployed task model with a decentralised identity model.

The recommendation is specifically the identity layer. ANP's DID mechanism carries no dependency on the rest of its stack. Its description and discovery protocols are a different matter: they overlap the A2A AgentCard and compete with it rather than stacking beneath it.

That distinction only became visible once the axes were scored separately. A single overall verdict per protocol would have hidden it.

What we are not claiming

Every score is one rater's reading. We have no inter-rater agreement statistic, and we say so in the abstract rather than burying it in a limitations paragraph.

What we offer instead is auditability. The scores live in a YAML file, with an evidence string and a citation per cell. Every table and figure is generated from that file, so a table cannot drift from the data behind it. You can disagree with a cell and point at the exact thing you disagree with.

Five of the thirteen protocols are surveyed but not scored. They are the commerce-settlement ones, where six of the ten axes have no referent. A settlement envelope has no task model and no multi-modal content handling. Scoring one of them while the others stayed unscored would move the averages in our own figures for reasons of scope rather than substance. UCP is the interesting case: it would return a defensible score, and we still left it out.

The specifications also move faster than a review cycle. So the survey declares a 2026-Q1 cutoff on page 1, and pins the exact version and commit behind every claim. That precision is not pedantry. One protocol we read ships two release tags that declare the same version string and differ by thousands of lines.

The loop we did not expect to find

We proposed three extensions against the gaps. A workspace naming scheme, a context-selection envelope, and an egress-attestation envelope. We expected three independent proposals.

They turned out to be three legs of one loop. A guardrails envelope with no normative workspace name has to invent a duplicate naming standard. A selection envelope with no boundary notion has nothing to key its policy on. A naming scheme with neither has scope but no enforcement.

Which is the survey's clearest argument for why the rubric axes cannot be worked on independently. Identity, authorisation, and security are one problem wearing three hats, in any deployment that crosses an organisational boundary.

All three are specification sketches for the next revision cycle, not validated standards, and we would rather label them that way than oversell them.

Read the full paper: https://hal.science/hal-05751265

If a cell is wrong, we would rather hear it than publish it. Several have already changed that way, including one that dropped a score after a lead author told us we had been too generous.

Every agent.
Accounted for.