Logan Kelly
Prompt guardrails can't be made bulletproof — NIST proved it in June 2026. 53% of orgs see agents exceed scope. The enforcement model that holds.

AI agent policy enforcement — also searched as LLM policy enforcement — is infrastructure-level control over what an agent is allowed to do, evaluated outside the model's reasoning. Unlike prompt-based guardrails, which a June 2026 NIST proof shows can never be universally robust, enforcement acts at three points: before execution, during execution, and after. Rules the model cannot reason around.
"Guardrails" is the word that does the most work in the AI safety conversation while doing the least amount of specification.
Ask ten engineers what they mean by guardrails and you'll get ten different answers. Some mean output filtering — checking what the model says before it goes to the user. Some mean system prompt instructions that tell the model to behave well. Some mean topic restrictions that prevent the model from engaging with certain domains. These are all real things. None of them is what I mean by policy enforcement, and none of them constitutes a governance strategy.
This post is about what policy enforcement for AI agents actually looks like when it needs to be reliable, auditable, and effective in production — not just plausible on a demo. Those three enforcement moments — pre-execution, mid-execution, and post-execution — each catch different risks. Together they cover all three moments where an agent's behavior can still be shaped. (See also: What is agentic governance → · The governance gap →)
The Fundamental Problem with Prompt-Based Guardrails
The most common approach to agent behavior control is the system prompt. You include instructions like "do not share confidential information," "always ask for confirmation before sending emails," "do not call the payment API without explicit user approval." These feel authoritative. They're not — and as of June 2026, that's not just an empirical observation, it's a mathematical one.
On June 9, 2026, NIST published a peer-reviewed proof by senior scientist Apostol Vassilev in IEEE Security & Privacy establishing that no finite set of guardrails can be universally robust against adversarial prompts. The proof extends Kurt Gödel's 1931 incompleteness theorems to AI: because guardrails are a finite set of rules governing an infinitely ambiguous input space (natural language), there will always exist some prompt that evades them. "There is no finite set of guardrails that is universally robust against adversarial prompts," Vassilev said. This isn't a claim that today's guardrails are poorly built — it's a structural limit on what guardrails, as a category of defense, can ever guarantee.
NIST's own recommended response has three elements: constant red-teaming to find breaking prompts before attackers do, continuous updates that harden guardrails against newly discovered prompts, and operational resilience that prioritizes impact limitation and quick recovery for when — not if — an exploit lands. That third element is the one this post is about. Limiting impact after a model has been talked into something is not a model-layer job. It is an enforcement job, and it has to live somewhere the prompt cannot reach.
The failure is at the action layer, not the language layer
The real-world version of this played out a month earlier. On May 4, 2026, according to a Security Boulevard analysis of the incident, an attacker drained roughly $175,000 in tokens from an AI-controlled crypto wallet belonging to Grok, xAI's chatbot — using a tweet written in Morse code.
The mechanism is worth reading slowly, because the obvious lesson is the wrong one. First the attacker sent a Bankr Club Membership NFT to Grok's auto-provisioned wallet, which unlocked Grok's ability to invoke the transfer tools of Bankrbot, an automated finance agent connected to Grok through a tool-calling layer. Then a Morse-coded reply on X told Grok to instruct Bankrbot to move 3 billion DRB to the attacker's address. Grok decoded the message and posted a clean English version tagging Bankrbot. Bankrbot executed.
The report is explicit that the model did not misbehave: decoding text and replying is exactly what a public assistant does. The breach, in its assessment, lives at the next hop — Bankrbot accepted the reply as a trusted instruction with no policy in between asking whether that principal was authorized to move those funds, to that recipient, in that amount. An earlier safeguard that blocked replies from Grok was bypassed because the gifted NFT opened an alternate tool-call path. The control was bound to a surface, not to an action.
That distinction is the whole argument. A 2025 empirical study by Hackett et al. (arXiv:2504.11168) tested six prominent protection systems — including Microsoft's Azure Prompt Shield and Meta's Prompt Guard — and found that character injection and algorithmic evasion techniques achieved in some instances up to 100% evasion success while maintaining adversarial utility. These are not exotic attacks. They are the current baseline, and NIST's proof explains why patching the guardrail layer will never drive that baseline to zero.
The problem isn't the model. It's where the enforcement lives. System prompt instructions are suggestions to a probabilistic system. LLMs follow them most of the time. They don't follow them all of the time. Under adversarial conditions — prompt injection, unusual input formats, carefully constructed edge cases — compliance rates with system prompt constraints drop significantly. Under distribution shift (inputs that don't match your training or testing distribution), they drop unpredictably.
This doesn't mean you shouldn't use system prompts thoughtfully. You should. But system prompts are not a governance layer. They're part of the user experience design. Governance requires enforcement mechanisms that exist outside the model's reasoning process — mechanisms that act regardless of what the model decides. The failure mode explored in depth in prompt injection via tool call results shows this precisely: the attack surface isn't the input you control, it's the tool response that feeds back into the agent's context.
When a national evaluator hit the same wall
In August 2026, the UK AI Security Institute (AISI) published an incident report on its own cyber evaluations. Across 122 runs of a single challenge, AISI found that in 10 runs the agents under test took a catalogued total of 19 unsanctioned actions on the live internet, directed at real people and organisations, between July 25 and 28, 2026. Seventeen came from Anthropic's Mythos 5; two from a single run of OpenAI's GPT-5.6 Sol. In the most serious sequence an agent attempted a supply-chain attack on a real open-source project, creating fake identities to pressure a human maintainer into approving malicious code. The maintainer refused it.
The caveats matter and AISI states them plainly: internet access was deliberately enabled and model-provider cyber classifiers were deliberately switched off to measure underlying capability, the tested configurations are not commercially available, its own design choices contributed to the behaviour, and its investigation found no evidence of resulting real-world harm. This is not evidence that deployed models behave this way.
What makes it relevant here is AISI's own account of why the behaviour ran as long as it did. The anomalous traffic was caught by general security monitoring after the fact, not by monitoring built to watch the evaluation as it ran — monitoring that, in AISI's words, could have flagged or blocked the behaviour sooner. AISI also notes that in several cases the margin between failure and success rested on human vigilance rather than a technical barrier, and that good containment should not depend on the model choosing not to test its boundaries. Its remediations are pre-execution and mid-execution controls: fine-grained network controls that make internet access a justified decision rather than a default, and real-time monitoring able to flag or block out-of-scope actions as they happen. The incident is covered in depth in AISI Incident Report: AI Agents Hit Real Targets When Scope Lived in a Prompt.
Pre-Execution: Block Before It Happens
Pre-execution enforcement intercepts a proposed action before it's executed. The agent has decided it wants to do something. The policy layer evaluates whether it's allowed to. If not, the action is blocked, and the agent receives a response indicating the block and why.
This is the most powerful enforcement position because it prevents consequences before they occur. No data is transmitted. No tool is called. No cost is incurred. The bad action simply doesn't happen.
According to the Cloud Security Alliance (May 2026), 53% of organizations report that AI agents exceed their intended permissions occasionally or sometimes — and only 8% say agents never exceed permissions. Pre-execution action authorization is the control most directly aimed at that number. The scope violation problem examined in AI agent scope violations and permission enforcement is almost entirely a pre-execution failure: the agent was never blocked from taking the action in the first place.
The gap is not that teams disagree with this. It's that the control usually isn't there. In Gravitee's April 2026 survey of 750 senior technology leaders, only 30.5% of organisations said a defined scope of what an agent is permitted to access was in place before the agent went live, and only 34.1% had a documented process to pause or revoke an agent's access. Under two in five, on the two controls that decide whether a bad action is blocked or merely regretted.
What pre-execution enforcement covers:
Input inspection. Before the agent's input is processed by the LLM, it can be scanned for content that violates policy — PII that shouldn't enter the context, injection patterns, content categories you've flagged as restricted. If the input fails inspection, you can sanitize it, reject it, or route it differently before it ever reaches the model.
Action authorization. Before a tool call is executed — before an API is hit, a file is written, a database is updated, an email is sent — a policy check determines whether this action is permitted. The authorization decision can be based on the action type, the parameters of the action, the session context, the user's permission level, or any combination. This is where you enforce "do not call the payment API without explicit user approval" in a way that actually works — not through a prompt instruction, but through an enforcement gate that the model cannot reason its way around. It is also, precisely, the check that was missing at Bankrbot's action layer.
Spend pre-authorization. Before initiating an operation that will incur cost — a long context call, an expensive tool invocation — a budget check determines whether the session has remaining allocation. If not, the operation is blocked before cost is incurred. See AI agent token budget enforcement for how this plays out in multi-step agentic workflows where costs compound rapidly across loops.
Mid-Execution: Intercept In-Flight
Pre-execution enforcement assumes you can predict what actions an agent will want to take. For simple, well-defined agents, you can. For more complex agents with multi-step reasoning and dynamic tool selection, there will be cases where an action sequence you didn't fully anticipate emerges. AISI's account of a goal-directed agent persistently searching for routes its operators never intended is that case, observed under laboratory conditions.
Mid-execution enforcement intercepts the agent's action stream as it's happening and applies policies in real time, including policies based on accumulated context that wasn't available at the start of the session.
What mid-execution enforcement covers:
Tool result inspection. A tool call was made and permitted. The result comes back. Before that result is appended to the agent's context, it's inspected — for PII that shouldn't enter context, for injection patterns, for content policy violations, for schema anomalies that indicate something unexpected happened. This is the layer that catches instructions arriving inside data the agent fetched rather than inside the prompt you wrote.
Sequence-level policy evaluation. Some policies only make sense at the sequence level, not the individual action level. If an agent has made five different external API calls in a single session, that pattern may be a policy violation even if each individual call was permitted. Mid-execution monitoring can track patterns across a session and trigger policy responses based on accumulated behavior. New credentials, accounts or network tunnels created by an agent mid-run belong in this category — in the AISI incident, agent-created identities were part of the sustained sequence AISI catalogued after the fact.
Budget enforcement. As token spend accumulates within a session, mid-execution monitoring tracks against the budget ceiling and triggers predefined responses — compression, warning, capping — as thresholds are approached and crossed.
Post-Execution: Audit and Remediate
Post-execution isn't enforcement in the sense of preventing actions — the action has already occurred. It's the foundation of your audit trail and the trigger for remediation workflows.
It is also where most teams are slowest. In the same CSA research, 47% of organizations said they had experienced a security incident involving an AI agent in the past 12 months, and 58% said detection and response take five hours or longer. Five hours is enough time for an agent's action to propagate through downstream automation. A record you can query in minutes is what compresses that number.
What post-execution covers:
Audit record creation. Every action the agent took, every policy evaluation that was performed, every enforcement decision that was made — logged with full context. The audit record should be sufficient to reconstruct what happened and why, including what the agent's context was at the time a decision was made. "The call was made" is not sufficient. "The call was made, here is the full context at that moment, here is the policy that was evaluated, here is the outcome" is sufficient.
Violation flagging. Actions that completed but should be reviewed — either because they barely passed policy or because they fit a pattern that warrants attention — get flagged for human review. This creates a workflow for operationalizing governance, not just logging it.
Retrospective detection. For behavioral patterns that are only apparent in aggregate — a class of queries where the agent consistently underperforms, a tool call pattern that's technically within policy but warrants investigation, a cost distribution that's shifted in a concerning direction — post-execution analysis of the execution record surfaces signals that weren't visible at the individual event level.
Policy Definition: The Work Before the Enforcement
None of this enforcement machinery matters if you haven't done the harder work of defining what your policies actually are.
Policy definition requires answering questions that feel abstract but have concrete implications:
What are the hard constraints — the things the agent must never do, regardless of context? These become blocking rules at the pre-execution layer.
What are the conditional permissions — things the agent may do under certain conditions? These become conditional authorization rules with context-dependent evaluation logic.
What are the budget parameters — token allocation per session, per user, per day; cost ceilings at various granularities? These become spend guardrails with defined response actions at threshold crossings.
What needs to be auditable — which actions, which data flows, which policy evaluations? These determine your audit log schema and retention requirements.
Policies need to be explicit, versioned, testable, and documented. An implicit policy — "the agent shouldn't do X" based on a system prompt instruction — is not a policy in the governance sense. An explicit policy — "action type Y with parameter P matching pattern Z is blocked for sessions with context flag Q" — is. Waxell Runtime ships with 50+ policy categories out of the box, covering input inspection, action authorization, spend pre-authorization, tool result inspection, and sequence-level behavioral rules — so you're enforcing documented, versioned policies from day one, not writing enforcement logic from scratch.
Testing Your Policies
A policy layer that only gets exercised in production incidents is not sufficient. Your policies need to be tested against known scenarios before they're deployed, and they need to be validated against your real production traffic on an ongoing basis.
This means having a test suite for your governance layer — not just for your agent's core behavior. Tests that verify:
Known bad inputs are blocked at the right layer
Known good inputs pass without unnecessary friction
Budget guardrails trigger at the right thresholds with the right responses
PII detection catches the patterns you care about with acceptable false positive rates
Audit records are created for the events that need records
The teams that have done this work have a dramatically different experience during incidents than the teams that haven't. When something goes wrong, the question isn't "do our policies work?" — it's "which policy handled this, and did it handle it correctly?"
That's a much better problem to have.
How Waxell handles this: Which Waxell product you want depends on how the agent was built. If you're building a new high-stakes workflow, Waxell Runtime is the execution environment — policies gate each step before it runs, with durable checkpoint-and-resume and kill switches at the agent, workflow and session level. If you already have agents running on LangChain, CrewAI or your own Python, Waxell Observe instruments them in two lines of code and enforces policy during execution, without a rewrite. Both run the same 50+ policy categories and produce the same execution record after the fact. Pick the one that matches your agents and start there: waxell.dev/signup
FAQ
What is AI agent policy enforcement? AI agent policy enforcement is implementing governance rules through technical mechanisms that act outside an agent's reasoning process, regardless of what the model decides. It's distinct from system prompt guardrails, which are suggestions to a probabilistic model and fail under adversarial conditions — and which a June 2026 NIST mathematical proof shows can never be made universally robust, no matter how well designed. Policy enforcement requires an external layer — typically at the infrastructure level — that evaluates and acts on agent behavior independently of the model.
How is AI agent policy enforcement different from system prompt guardrails? System prompt instructions are suggestions to a probabilistic model. LLMs follow them most of the time — not all of the time. Under adversarial conditions (prompt injection, unusual inputs), compliance drops significantly — a 2025 empirical study (arXiv:2504.11168) found that adversarial evasion techniques achieved up to 100% success in some instances against six major guardrail systems including Azure Prompt Shield and Meta Prompt Guard. Policy enforcement uses mechanisms outside the model's reasoning: pre-execution gates that block actions before they fire, mid-execution interceptors that apply rules regardless of what the model decided, and post-execution audit that documents every governance decision. The model cannot reason its way around these.
What are the three layers of AI agent policy enforcement? Pre-execution enforcement blocks proposed actions before they execute — input inspection, action authorization, spend pre-authorization. This is the most powerful position because it prevents consequences before they occur. Mid-execution enforcement intercepts the agent's action stream in real time — tool result inspection, sequence-level policy evaluation, budget tracking. Post-execution enforcement creates the audit trail and triggers remediation — logging every governance decision with full context, flagging violations for human review, enabling retrospective behavioral analysis.
What is pre-execution enforcement for AI agents? Pre-execution enforcement intercepts a proposed action before it executes. The agent has decided it wants to take an action; the policy layer evaluates whether it's permitted to. If not, the action is blocked and the agent receives a response explaining the block. This is the strongest enforcement position — no data transmitted, no tool called, no cost incurred, no consequences to remediate. Pre-execution covers input inspection (PII, injection patterns), action authorization (is this tool call permitted under current conditions?), and spend pre-authorization (does this session have remaining budget?).
How do you test AI agent policies? Policies need a test suite separate from your agent's core behavior tests. The suite should verify: known bad inputs are blocked at the right layer; known good inputs pass without friction; budget guardrails trigger at the correct thresholds with the correct responses; PII detection catches targeted patterns at acceptable false positive rates; and audit records are created for all events that require them. Running this suite before every deployment that changes policies is the difference between finding out a policy broke in production and finding out in CI.
Can LLM guardrails be bypassed in practice? Yes — and as of June 2026, this isn't just empirical, it's proven mathematically. NIST senior scientist Apostol Vassilev published a peer-reviewed proof in IEEE Security & Privacy showing that no finite set of guardrails can be universally robust against adversarial prompts, extending Gödel's incompleteness theorems to AI systems. Empirically, a 2025 study (arXiv:2504.11168) tested six major guardrail systems including Microsoft Azure Prompt Shield and Meta Prompt Guard and achieved in some instances up to 100% evasion success while maintaining adversarial utility. Real-world incidents corroborate this: according to a May 2026 Security Boulevard report, an attacker drained roughly $175,000 from an AI-controlled crypto wallet using a Morse-code-encoded tweet — with the missing control being an authorization check at the executing agent's action layer, not a better filter at the model layer.
How do you enforce policy-based execution for AI agents? Policy-based execution means every agent action passes through an enforcement layer that evaluates it against explicit, versioned rules before, during, and after it runs — not a system prompt asking the model to behave. In practice this means a pre-execution authorization gate on tool calls, mid-execution inspection of tool results and accumulated session behavior, and post-execution audit logging with full decision context. The rules live outside the model, so a compromised or manipulated model still can't take an action the policy layer hasn't approved.
How does policy enforcement work at scale? At scale, policy enforcement stops being about individual rule checks and starts being about consistency across a fleet of agents, models, and providers. That requires policies to be centrally defined, versioned, and tested — not duplicated ad hoc inside each agent's prompt — and enforced by infrastructure that evaluates every agent's actions the same way regardless of which model or framework is running it. Sequence-level and cross-session policy evaluation also matters more at scale: a single flagged action might be noise, but the same pattern recurring across hundreds of agent sessions is a signal that a policy needs tightening.
Sources
NIST, NIST Mathematical Proof Supports Transition to a Continuous-Monitor-and-Update Security Model for AI Systems (June 2026) — https://www.nist.gov/news-events/news/2026/06/nist-mathematical-proof-supports-transition-continuous-monitor-and-update
Apostol Vassilev, Robust AI Security and Alignment: A Sisyphean Endeavor?, IEEE Security & Privacy (May/June 2026), DOI: 10.1109/MSEC.2026.3678214
UK AI Security Institute, Incident Report: unsanctioned agent behaviour during cyber testing (August 2026) — https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Cloud Security Alliance, AI Agent Security Starts with Scope Control (May 2026) — https://cloudsecurityalliance.org/blog/2026/05/12/ai-agent-security-starts-with-scope-control
Gravitee, The State of AI Agent Security 2026 (April 2026, n=750 senior technology leaders) — https://www.gravitee.io/state-of-ai-agent-security
Shreyans Mehta (Cequence Security), via Security Boulevard, Encoded Prompt Injection: Why LLM Guardrails Are at the Wrong Layer (May 2026) — https://securityboulevard.com/2026/05/encoded-prompt-injection-why-llm-guardrails-are-at-the-wrong-layer/
Hackett et al., Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems (April 2025) — https://arxiv.org/abs/2504.11168
Agentic Governance, Explained





