Logan Kelly

What Is an AI Control Plane?

What Is an AI Control Plane?

An AI control plane governs what agents can do, before they act — not after. See how Waxell's five products enforce it.

Waxell blog cover: diagram of an AI control plane sitting above agents, models, and tools

AI agent cost control is the practice of enforcing token budgets, session limits and spend ceilings at the infrastructure layer, so a run stops before its next model call. Cost observability reports what an agent spent once the spend has happened. Cost control decides whether the next call happens — the last point at which the outcome can still change.

The distinction is not academic. In September 2026, The Futurum Group surveyed 1,636 enterprise technology decision makers and found that 46.9% of organizations are spending more on AI than they planned. What they do about it is the interesting part: of the 767 organizations running over plan, 47.6% ask for more budget and 43.3% absorb the overrun and settle it later. Only 17.2% pause or reduce the initiative. "When nearly half of organizations are running AI spend over budget and only one in six responds by slowing the initiative down, the budget has stopped being a control," said Mitch Ashley, Futurum's VP and Practice Lead for CIO and Technology Buyers.

A budget nobody enforces is a forecast. And agents are unusually good at breaking forecasts.

The most widely circulated illustration is an engineer's own published account. Writing in Towards AI in October 2025, Kusireddy described deploying four LangChain agents coordinating over the Agent-to-Agent (A2A) protocol to research market data. The costs went $127 in week one, $891 in week two, $6,240 in week three, $18,400 in week four. Two of the agents had got stuck in a conversation loop with each other — "For 11 days," as the account puts it. "While we slept. While we worked. While we believed 'it's just running smoothly.'" Total: $47,000, and it ended because a human eventually pulled the plug. (This is a single first-person account rather than an audited post-mortem; the company is not named. We cover the incident, and what a budget ceiling would have changed, in The $47,000 Agent Loop →.)

There's a mental model mismatch that makes this kind of disaster predictable, and it comes down to this: engineers think in requests. Agents run in loops.

When you're building a traditional API-backed feature, cost math is simple. One user action equals one API call. One API call has a known cost. You multiply by your DAU and you have your monthly bill, approximately. Budget accordingly.

When you're building agents, this model falls apart. An agent doesn't make one call. It reasons, retrieves, calls tools, observes the results, reasons again, calls more tools, synthesizes — and each of those steps is a context window, and each context window grows as the conversation goes on because everything that came before gets appended. The token count isn't static. It compounds. (For a deeper look at how context accumulation specifically drives spend, see The Real Cost of AI Agent Context Windows →.)

Agent costs don't scale linearly with requests: they compound with context accumulation, loop behavior, and tool call overhead, which is what makes proactive enforcement necessary rather than merely tidy. A 5-step agent loop serving 200 concurrent users can cost 10× what naive per-request estimates suggest.

Most teams have no structural defence against that. A Gartner survey of 353 data & analytics and AI leaders, fielded from November through December 2025 and published in March 2026, found that only 44% of organizations have adopted financial guardrails or AI FinOps practices. The majority are flying blind until the bill arrives.

Why Do AI Agent Costs Multiply With Each Loop?

Take a modest agent. It handles customer support queries. For a typical question, it goes through this cycle: initial reasoning call, two tool calls (retrieve relevant docs, look up account status), a synthesis call to compose the response. Four LLM calls. Say your context window at each step is, on average, 4,000 tokens. That's 16,000 tokens for one resolved query.

That doesn't sound terrible. You're building a support product. You're charging for it. Fine.

Now add scale. You're serving 200 concurrent users. You're now at 3,200,000 tokens per "round" of queries. If an average session involves five exchanges before it's resolved, you're at 16,000,000 tokens per hour of support traffic.

Now add the thing nobody budgets for: variance. Some queries don't resolve in five exchanges. Some agents get into loops — they call a tool, the tool returns something unexpected, they reason about it, call the tool again, get a similar result, reason again. Without a hard stop, a single looping session can consume 100× the token budget of a normal one. You don't notice it until it's in the tail of your cost distribution, and by the time it's notable in your aggregate, it's happened hundreds of times.

This is not a hypothetical, and it is not confined to startups. Uber's CTO disclosed in April 2026 that the company had spent its entire annual AI budget in four months; TechCrunch, citing Bloomberg, reported in June that Uber's response was a hard monthly cap of $1,500 per employee per agentic coding tool, with exceptions granted by permission. Worth noting what the same reporting says preceded the burn: Uber had encouraged staff to use AI “as much as possible” and ranked usage competitively on internal leaderboards. The visibility was not the control. The cap was.

Where Do AI Agent Cost Spirals Actually Come From?

Context accumulation. The longer a conversation runs, the more tokens get included in every subsequent call. A 20-turn conversation doesn't cost 20× a 1-turn conversation — it costs significantly more, because every turn includes the entire history. If you're not actively managing context window size (through summarization, pruning, or hard turn limits), you're letting cost compound with every exchange.

Tool call overhead. Each tool call requires the agent to explain what it's calling and why, get the result back, and reason about the result — all of which adds tokens. Agents that call tools aggressively, or that call tools whose results are verbose, pay a significant overhead per call. A tool that returns 500 tokens when 50 would do is costing you 10× more than necessary.

Retries. When a tool call fails, or when the model's output doesn't pass a validation step, the agent retries. Retries are full calls at full cost. An agent that retries aggressively on flaky tools can run up a large bill in a short time.

Concurrent sessions at peak. If your agent is customer-facing, you have peak hours. During peak, the number of concurrent sessions multiplies your per-session costs. If your agent also tends to take longer on high-complexity queries, and high-complexity queries cluster at peak, you can see where this goes.

Pricing models that move underneath you. On June 1, 2026, GitHub switched Copilot from a flat request-based rate to token-usage billing. Ahead of the change, TechCrunch collected developers projecting sharp increases — one claiming a roughly $29 monthly bill would rise to nearly $750, another sharing a screenshot that appeared to show a jump from about $50 to some $3,000 — while other Copilot users argued the increases were avoidable with more careful usage. Whichever camp is right, the structural point stands: the cost of a given agent workflow is not a constant you can budget once. (We covered that migration in Copilot Billing Shock →.)

What Are the Most Effective Strategies for Controlling AI Agent Costs?

1. Session-level token budgets with hard stops. The most direct intervention. Set a maximum token budget per session. When a session reaches the limit, execution stops — cleanly, and with the event recorded. The key word is "hard stop": a soft warning the model can reason around is not a control. A budget enforcement layer has to be able to stop a run before it burns past the ceiling, not notice afterwards that it did.

This sounds aggressive. In practice, a well-tuned budget catches the outlier sessions — the loops, the unusually complex queries — without touching the median session at all. The $47,000 loop ran for eleven days because the only thing that could stop it was a person noticing. For a fuller breakdown of why alerts aren't enforcement, see The $400M AI FinOps Gap →.

2. Context pruning and compression. Rather than including the full conversation history in every call, summarize older turns. After every N exchanges, run a lightweight summarization call that compresses the history, replace the raw history with the summary, and proceed. You lose some fine-grained context but retain the semantic substance. For most agent tasks, that trade is favorable, and the cost of the summarization call is paid back immediately in the reduced token count of every subsequent call.

3. Tiered model routing. Not every step in your agent's reasoning requires your most capable — and most expensive — model. Tool selection, parameter extraction, routine classification: these can often be handled by smaller, faster, cheaper models. Reserve frontier models for the steps that genuinely need them. The cost difference on simple tasks is often 10× or more.

This one has a limit worth naming. Gartner's Nitish Tyagi, writing in June 2026 about AI coding agents, put it this way: "Token discipline will not emerge through developer choice alone, as developers tend to optimize for speed and convenience over cost efficiency." Routing is an architecture decision, not a habit — if it depends on someone remembering, it will drift.

4. Real-time spend visibility — and then something that acts on it. You do need to know about a spiraling session while it is happening. Track agent costs in real time, alert at 60% of the session ceiling, and set a fleet-level alert when aggregate spend for the hour is trending past your daily allocation.

But visibility is the input to a control, not the control itself. An independent survey of 500 finance leaders at large US and UK enterprises, conducted by Sapio Research in February 2026 and published by DoiT in June, found 79% had experienced AI cost overruns in the preceding twelve months — and that the organizations self-assessing as most FinOps-mature reported the highest overrun rates. DoiT's own reading of that is the honest one, and it is worth quoting rather than paraphrasing: "Maturity surfaces problems. It does not prevent them, and presenting it as prevention sets up a promise the data will not support." (DoiT attributes the mature cohort's higher rate partly to running larger programs and partly to simply detecting overruns that less instrumented organizations never catch — so this is not an argument against instrumentation. It is an argument that instrumentation is not the same thing as enforcement.)

What Do AI Agent Budget Guardrails Look Like in Practice?

A budget guardrail isn't a feature you add to your agent code. It's a policy evaluated outside it — something that knows the spend to date, compares it against a defined ceiling, and returns a decision the agent cannot argue with.

The decision has to arrive at the right moment. A pre-execution check stops a run that should never have started. A mid-execution check, run between steps, is what catches budget exhaustion inside a long-running agent — the exact failure mode in the eleven-day loop, where the first call was perfectly reasonable and the four hundredth was not. An enforcement layer that only evaluates at the start of a run will watch a loop spend $47,000 without ever being consulted again.

Which response is right depends on your product. A customer support agent might warn first and throttle before it blocks, because a hard stop mid-conversation creates a bad experience. An internal automation agent can be blocked outright — the job can be requeued when the budget resets. What matters is that the choice is made at policy definition time and applied consistently, rather than improvised during an incident. This is the same argument we make about governance generally: a control that lives in a document instead of the execution path is not a control, which is the point of agentic AI governance on the execution path →. (See also: What is agentic governance →)

Who Owns AI Agent Cost Governance?

Cost governance for agents needs cross-functional alignment that most organizations haven't established yet.

Engineering sets the technical ceiling. Finance needs to understand that agent costs are variable in ways traditional SaaS infrastructure costs are not. Product needs to understand the cost implications of features that increase session depth or tool call frequency. Without that alignment you get a spike, a panicked response, and a retroactive patch — instead of a policy that was in place beforehand.

Futurum's data says something uncomfortable about where that conversation currently lands: when organizations go over plan, asking for more budget and absorbing the overrun are together more than five times as common as slowing the initiative down. The time to build the alignment is before the spike. It's a boring conversation. Have it anyway. If you're thinking about how this fits into a broader operational discipline, AgentOps: The Discipline Missing From Your AI Deployment Stack → covers the wider operational stack.

Agent costs are controllable. The math isn't mysterious and the tooling exists. What's usually missing is a policy layer that turns a number someone agreed to into a decision made in the execution path, automatically, without requiring an engineer to be watching.

How Waxell handles this: Cost is one of the 50+ policy categories Waxell Observe ships with, and budget limits are defined in the governance plane rather than in agent code — per agent, per model, or per time period. Waxell's Budgets capability evaluates those limits before a run begins and stops execution when a limit is reached, recording the enforcement event with the agent, the model, the limit and the timestamp. For long-running agents, Observe can also check policy between steps, which is what catches budget exhaustion partway through a run rather than only at the start. A policy check returns a decision the agent code does not get to override; warn, throttle and block are among the documented responses, so you can tune the outcome to the workload. Limits can be tightened or relaxed without a deploy, because they live outside the agents they govern. Real-time spend telemetry gives per-session token visibility across the agents you have instrumented, and instrumenting them is two lines of Python against the 200+ Python libraries Observe auto-instruments — no rebuild. Start free with Waxell Observe →: pip install waxell, 10,000 traced executions a month on the free plan.

Frequently Asked Questions

Why do AI agent costs spiral in production? Agent costs spiral because they compound rather than scale linearly. Each turn in a conversation extends the context window, making every subsequent LLM call more expensive. Tool calls add overhead at each step. Loops — where agents retry a failing operation or get stuck responding to each other — can run up 100× the normal session cost. Without a policy layer enforcing a budget ceiling, one outlier session can consume as much as hundreds of typical ones. The most-cited example is an engineer's published account of four agents coordinating over A2A, two of which looped for eleven days and cost $47,000 before a human intervened.

How do we prevent a single AI agent from consuming runaway LLM tokens? Put a per-agent ceiling outside the agent and evaluate it in the execution path, not in the prompt. Three things make it work: a spend or token limit scoped to that agent rather than to the account, a check that runs between steps as well as before the run starts, and a response the agent cannot talk its way past. A limit expressed as an instruction in a system prompt is a suggestion; a limit evaluated by a policy engine before the next call is a control. Rate limits on invocation frequency are a useful second layer for agents that fail by retrying rather than by reasoning.

How do you calculate the real cost of an AI agent session? Multiply the average number of LLM calls per session by the average context window size at each call. For a typical 4-step agent handling a query with a 4,000-token average context per step, that's 16,000 tokens per resolved query. At 200 concurrent users and 5 exchanges per session, you're at roughly 16 million tokens per hour of traffic — before accounting for variance from long sessions or loops. Most teams' initial estimates are off by 5–10× before they do this math explicitly.

What is a token budget for AI agents? A token budget is a hard limit on the total tokens a single agent session is allowed to consume. When the session reaches the limit, the policy layer returns a predefined decision: warn and continue, throttle, or stop the run. A well-tuned budget catches outlier sessions — loops, unusually long conversations — without affecting typical sessions at all, which is why a ceiling set near the 99th percentile of normal usage rarely costs you anything in practice.

What causes AI agent cost explosions? Four mechanisms cause most of them: context accumulation (the full conversation history included in every subsequent call), tool call overhead (verbose tool responses inflating every context window), retry behavior (failed calls at full cost), and concurrent session peaks (multiplied per-session costs during high traffic). Any one is manageable. When they compound — a peak traffic moment with verbose tools and some agents in retry loops — costs can spike dramatically in minutes.

How do you stop an AI agent from going over budget? Budget enforcement has to sit outside the agent, not in its system prompt. The policy layer compares spend to date against the defined budget and returns a decision — not a suggestion the model can override. Effective implementations alert well before the ceiling while there is still time to act, compress or prune context as the session lengthens, and throttle or block at the ceiling. The key word is "hard stop": a soft warning the model can reason around does not constitute cost governance.

What's the difference between AI agent cost monitoring and cost enforcement? Cost monitoring tells you what was spent — dashboards, per-session breakdowns, alerts when thresholds are approached. It is asynchronous: by the time an alert fires, the spend has happened. Cost enforcement evaluates the next call against a ceiling before it goes out, and stops the run if the ceiling has been reached. Uber is the cleanest illustration of the gap: its AI usage was visible enough to rank on internal leaderboards, and the company still spent its annual AI budget in four months. The fix was a hard per-employee cap.

Why did AI coding costs become a budget problem in 2026? Because the pricing model changed underneath the workflow. Vendors moved from seat-based licensing toward consumption-based pricing — GitHub switched Copilot to token-usage billing on June 1, 2026 — which makes spend a function of how hard an agent works rather than how many people are licensed. Gartner predicted in June 2026 that AI coding costs will overtake the average developer's salary by 2028 on rising token consumption, and warned that many vendors lack transparency into how token consumption is calculated and billed, limiting an enterprise's ability to forecast it. Under consumption pricing, an unbounded agent is an unbounded invoice.

Sources

  • The Futurum Group, 46.9% of Enterprises Report AI Spend Over Budget in 2H 2026 (September 2026) — https://futurumgroup.com/press-release/46-9-of-enterprises-report-ai-spend-over-budget-in-2h-2026/ (2H 2026 CIO & Technology Buyers Decision Maker Survey, n=1,636; 767 organizations over plan)

  • Kusireddy, We Spent $47,000 Running AI Agents in Production. Here's What Nobody Tells You About A2A and MCP, Towards AI (October 16, 2025) — https://pub.towardsai.net/we-spent-47-000-running-ai-agents-in-production-heres-what-nobody-tells-you-about-a2a-and-mcp-5f845848de33 (first-person account from an unnamed engineering team; not independently audited. The Medium copy previously cited in this post was deleted by its author and now returns HTTP 410.)

  • Tech Startups, AI Agents Horror Stories: How a $47,000 AI Agent Failure Exposed the Hype and Hidden Risks of Multi-Agent Systems (November 14, 2025) — https://techstartups.com/2025/11/14/ai-agents-horror-stories-how-a-47000-failure-exposed-the-hype-and-hidden-risks-of-multi-agent-systems/ (secondary coverage; describes the setup as "four LangChain-style agents")

  • DoiT / Sapio Research, Why 79% of Enterprises Overspent on AI in 2026 (June 9, 2026) — https://www.doit.com/blog/ai-spending-survey (n=500 finance leaders, US and UK, organizations of 1,000+ employees; fielded February 2026)

  • Gartner, Gartner Identifies Three Pillars for Deriving Value from AI (March 9, 2026) — https://www.gartner.com/en/newsroom/press-releases/2026-03-09-gartner-identifies-three-pillars-for-deriving-value-from-ai (survey of 353 D&A and AI leaders fielded November–December 2025; 44% financial-guardrails adoption)

  • Gartner, Gartner Predicts AI Coding Costs Will Surpass Average Developer's Salary by 2028 as Token Consumption Surges (June 24, 2026) — https://www.gartner.com/en/newsroom/press-releases/2026-06-24-gartner-predicts-ai-coding-costs-will-surpass-average-developer-salary-by-2028-as-token-consumption-surges

  • TechCrunch, Uber caps employee AI spending after blowing through budget in 4 months (June 2, 2026) — https://techcrunch.com/2026/06/02/uber-caps-employee-ai-spending-after-blowing-through-budget-in-four-months/ (TechCrunch citing Bloomberg for the $1,500 cap; the four-month budget exhaustion was disclosed by Uber's CTO in April 2026)

  • TechCrunch, 'What a joke': GitHub Copilot's new token-based billing spurs consternation among devs (May 30, 2026) — https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/ (developer cost figures are user-reported projections and screenshots collected before the June 1 change took effect)

  • Waxell, Policy & Governance (Waxell Observe docs) — https://waxell.ai/docs/observe/features/governance (policy actions and mid-execution checks)

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.