Logan Kelly
In multi-agent systems, authority follows the handoff. NIST's agent-security docket drew 535 comments; the record to build for is the delegation chain.

When NIST's Center for AI Standards and Innovation opened its agent-security docket in January 2026, it asked a question most agent tooling still answers by accident. Question 1(e) of the Request for Information reads: "What unique security threats, risks, or vulnerabilities currently affect multi-agent systems, distinct from those affecting singular AI agent systems?" The docket, NIST-2025-0035, closed on 9 March 2026 with 535 published comments.
A delegation chain is the ordered record of which agent authorised which other agent to act, with what scope, on whose behalf. In a multi-agent system it is the smallest unit that explains an outcome, because the outcome is produced by a sequence of handoffs rather than by any one agent. It is also the unit least likely to exist, because a handoff happens between two agents and belongs to the instrumentation of neither.
That last property is the structural problem, worth stating plainly before the frameworks arrive: in a multi-agent system, authority moves further than the record of it. Each agent logs what it did. The event that matters — one agent conferring authority on another — sits in the gap between two logs.
What did the responses actually converge on?
On 23 September 2026 the Cloud Security Alliance published an editorial synthesis by Chandra Inguva, drawn from what he describes as "a purposive set of 100 public responses" to the RFI — roughly a fifth of the docket, selected rather than sampled. It is one practitioner's reading, not a NIST finding.
Two of its five lessons are about the handoff. Lesson 2, titled "Delegation Is a Trust Boundary," puts the mechanism in a single conditional: "If authority becomes transitive by default, privilege can expand beyond the user's original intent." Lesson 4, "The Trajectory Is the Audit Record," puts the consequence in another: "the final answer tells only a fraction of the story."
Those two sentences describe the same defect from opposite ends. Authority is transitive by default because delegation is implemented as a function call: the parent passes a task down, and whatever credentials, scopes and tool access the parent holds are simply there in the child's environment unless something deliberately narrows them. Nothing has to go wrong for privilege to widen. The default does it.
And the artifact that would show you the widening is the one that belongs to neither agent. A coordinator's log says it dispatched a task. A worker's log says it ran one. Neither record carries the thing a reviewer needs: what the coordinator was entitled to confer, what it actually conferred, and whether the worker could confer it again.
Why doesn't more instrumentation fix this?
Because the scope of instrumentation is the process, and the scope of the problem is the boundary between processes.
Tracing is the right tool for the inside of an agent. If you can install a library into the running process, you get the call graph: spans, parent-child relationships, tool calls with inputs and timings. That is exactly what Waxell Observe does — it traces the execution tree across a coordinator, its planners and their tool calls, linked by session and lineage, with parent-child relationships detected automatically and session IDs propagating through nested calls without manual wiring. It auto-instruments 200+ Python libraries, and the OpenTelemetry span tree it produces is a genuine trajectory.
The precondition is the whole story: it is Python, and it is a process you can get code into.
Production multi-agent systems increasingly are not that. The coordinator is your code; the thing it hands work to is Claude Code, Cursor, a vendor's agent behind an API, or a teammate's assistant. You cannot install a library into any of those. When a handoff crosses that line the span tree ends — not because tracing failed, but because it was never entitled to that side of the boundary. You are left with two trajectories that each look complete, and no join between them.
This is the same shape as the trust cascade problem: a compromised or merely mistaken upstream agent is accepted downstream because downstream has no independent view of what upstream was allowed to do. We covered the propagation half in how trust attacks cascade through multi-agent systems. The authority half is what the NIST docket keeps surfacing.
Is anyone standardising the delegation chain?
The work is under way and it is early.
NIST's National Cybersecurity Center of Excellence has a concept paper, Accelerating the Adoption of Software and AI Agent Identity and Authorization, applying identity standards to AI agents. The project page describes the NCCoE as "seeking feedback to help determine the scope, feasibility, and potential value of the project" — deciding whether to write a project description, not publishing one. The concept paper’s comment deadline was 2 April 2026.
Meanwhile the RFI itself treated agent-to-agent interaction as a secondary topic in the guidance it gave respondents. NIST listed nine questions for people with limited bandwidth to prioritise — 1(a), 1(d), 2(a), 2(e), 3(a), 3(b), 4(a), 4(b) and 4(d). The multi-agent question, 1(e), is not among them, and neither is 4(c), which covers interactions with counterparties including — as its fifth and final item — "other AI agent systems." That describes a prioritisation list, not what NIST believes; the questions were asked. But if you are waiting for a standard to tell you how to represent a delegation chain, the ordering is worth knowing.
The property arrives before the standard. Build for "the delegation chain can be reconstructed after the fact" now and adopt whatever schema eventually wins later; building for the schema first is the part that gets redone.
What does a handoff record have to carry?
Four things, and they are the four that a per-agent log structurally cannot supply on its own.
Who conferred what. Not that a task moved, but which authority moved with it — scopes, tool access, data reach. "Agent B started" is not a delegation record. "Agent A, acting for user U, conferred scopes X and Y on agent B" is.
Whether it could be conferred again. Transitivity turns one over-broad grant into an unbounded one. A record that stops at the first hop cannot tell you how far the grant travelled.
On whose behalf. The delegating human has to survive the hop. A chain that resolves to a service account at step two cannot answer the question an incident review asks first.
Where the chain was interrupted. A human approval, a policy decision, a refusal — these are the evidence that a control operated. If interventions live in a different system from the handoffs they interrupted, you can prove neither the order of events nor their relationship.
None of that is exotic. It is bookkeeping that has to be done by something standing between the agents rather than by the agents themselves.
How Waxell handles this
Waxell Connect is the surface the handoff crosses. Agents your team already runs — Claude Code, Cowork, Cursor, and others — join a shared workspace without SDK changes, and when one finishes a task Connect routes it to the next. Because the routing goes through Connect rather than between the agents privately, the handoff becomes an event that something other than the participants records: Connect logs the handoffs that cross it, along with file versions and agent actions, so you can see what ran, what changed, and which agent or person did it. That record is versioned, and it serves as an identity and audit record for work moving between agents no SDK can reach.
For the agents you can instrument, Observe supplies the inside of the picture: the parent-child span tree within a Python process, and policies evaluated during execution — before a step, between steps, after completion — returning a structured retry, escalate or halt rather than a dashboard alert after the fact. Delegation is one of the 50+ policy categories, alongside Identity, Control and Audit.
The split is not a packaging decision — it follows from instrumentability. Observe sees inside a process you can run code in; Connect sees across a boundary you cannot. A system spanning both needs both, and knowing which side of the line each agent sits on is the first inventory task, not a detail.
One honest limit, and it bears on the completeness property this post is about: Connect records the handoffs that route through it. Two agents that hand off directly — sharing a credential, calling each other's APIs, or coordinating through a channel Connect is not part of — produce no Connect record, and the audit trail will not tell you they happened. Coverage here is configured, not given. If your compliance story rests on the delegation chain being complete, the inventory of which handoffs actually traverse the coordination layer matters as much as the layer does.
FAQ
What is a delegation chain in a multi-agent system?
It is the ordered record of which agent authorised which other agent to act, with what scope, and on whose behalf. It differs from a trace in what it captures: a trace records that a call happened and how long it took; a delegation chain records what authority moved with the call. A system can have complete traces and no reconstructable delegation chain.
Why is delegation described as a trust boundary?
Because crossing it changes who can do what. The Cloud Security Alliance synthesis of NIST RFI responses frames it as a conditional — if authority becomes transitive by default, privilege can expand beyond the user's original intent — and the default does the damage. Treating delegation as a boundary means making authority explicit at each hop, preserving least privilege across it, and keeping enough provenance to reconstruct the chain afterward.
Are OpenTelemetry traces enough for multi-agent governance?
They are a strong input and an incomplete answer. Traces are emitted by the instrumented process, so they end where your instrumentation's reach ends — in a multi-agent system, usually at the first handoff to an agent you did not build. They also key on sessions and spans rather than on the delegating user and the authority conferred. The gap to close is attribution across the boundary, not trace volume inside it.
Has NIST published a standard for agent identity and delegation?
Not yet. The NCCoE concept paper Accelerating the Adoption of Software and AI Agent Identity and Authorization is at the feedback stage, with its project page describing the effort as determining the scope, feasibility and potential value of a possible project. The concept paper’s comment deadline was 2 April 2026. The separate CAISI Request for Information on agent security closed on 9 March 2026 with 535 published comments on docket NIST-2025-0035.
What should we do before any of this settles?
Three things survive whichever schema wins. Make the delegating human's identity ride the handoff rather than resolving to a service account at the first hop. Generate the handoff record somewhere other than the two agents involved in it. And inventory which agent-to-agent handoffs actually pass through a surface that can record them — that number is usually lower than teams expect, and it bounds everything else.
Does this only matter for large agent fleets?
No. Transitive authority is a property of the second hop, not of scale. A coordinator that spawns one worker which calls one vendor agent already has a three-link chain, two boundaries, and no single log that spans them.
Sources
National Institute of Standards and Technology, "Request for Information Regarding Security Considerations for Artificial Intelligence Agents", Federal Register, 91 FR 698, 8 January 2026
Regulations.gov, "Docket NIST-2025-0035 — public comments", National Institute of Standards and Technology
Cloud Security Alliance (Chandra Inguva), "Lessons Learned on Securing Multi-Agent Systems: NIST Agent Security RFI", 23 September 2026
National Cybersecurity Center of Excellence, "Software and AI Agent Identity and Authorization", NIST
National Institute of Standards and Technology, "AI Agent Standards Initiative"
The agents in a multi-agent system each keep an honest record of their own work. That is the problem. The event that decides what the system was allowed to do happens between them, and it appears in neither log unless something else is standing there to write it down.
Start free with Waxell Observe and one governed MCP upstream — 10,000 traced executions a month and 2 seats: https://waxell.dev/signup
Agentic Governance, Explained





