Logan Kelly

Agent Testing: The Same Model Scored 42% and 95%, and Only the Configuration Changed

Agent Testing: The Same Model Scored 42% and 95%, and Only the Configuration Changed

Agent testing measures a configuration, not a model. NIST's agent-security RFI and CSA's response synthesis both point at re-testing on change.

Waxell blog cover: the same model scoring 42% and then 95%

Anthropic reports that Claude Opus 4.5 scored 42% on CORE-Bench, and that one of its researchers then opened the benchmark and found rigid grading — 96.12 marked wrong where 96.124991… was expected — along with ambiguous task specifications and stochastic tasks that could not be reproduced exactly. On Anthropic's account, once those were fixed and the model was given a less constrained scaffold, it scored 95%.

Fifty-three points moved. The model was the same in both runs; the grading and the scaffold were not. It is the clearest available demonstration of what an agent test result actually describes.

An agent test result certifies a configuration, not an agent. The score belongs to a specific arrangement of model, prompt, tools, permissions, memory, topology, grading logic and harness — and it stops being evidence the moment any of those change. Treating a passing suite as a property of the agent, rather than a property of one configuration at one moment, is the assumption that puts stale assurance behind live systems.

What is the system actually under test?

The word "agent" hides the problem. When a team says it evaluated its agent, the thing that ran was a model inside a scaffold, wired to a tool set, holding a permission set, carrying memory, arranged in a topology, graded by logic someone wrote. The model is one component among several, and it is frequently the component that did not fail.

A practitioner writing on Hacker News described exactly this after trying to evaluate an agent with a benchmark-style suite. Broken URLs in tool calls dropped the score to 22. An agent calling localhost in a cloud environment got stuck at 46. Real CVEs were flagged as hallucinations by the grader. A missing API key in production produced a silent failure. Every run surfaced a real bug, and in the poster's account most of them came from system-level problems rather than model quality — the failures looked "more like software bugs than LLM mistakes."

This is not a reason to distrust evals. It is a reason to be precise about their subject. Anthropic's guidance says the same thing from the other direction: it treats it as essential that the agent in the eval behaves roughly as the production agent does, and that the environment introduces no further noise. The eval measures the whole assembly. Change the assembly and you have measured something else.

Why does a test result expire?

On 23 September 2026 the Cloud Security Alliance published a synthesis drawn from a purposive set of 100 public responses to NIST's Request for Information on security considerations for AI agents — the CAISI docket opened on 8 January 2026 at 91 FR 698, comments closing 9 March. Its author distilled five lessons from responses by cloud providers, security companies, model developers, enterprise software firms and researchers. CSA calls the fifth the strongest cross-cutting lesson of the set, and it bears directly on testing:

Assurance belongs to the system, not the model. Security depends on the configured agent system — its tools, permissions, data boundaries, memory, orchestration logic, network access, approval gates and surrounding controls. Organisations therefore need production-realistic environments, and because the system keeps evolving, "assurance should be continuously refreshed rather than treated as a one-time gate."

Six axes are named there: models, prompts, tools, permissions, memory, topology. Each moves independently of the others, none of them is the eval suite, and a change to any one can invalidate a result the suite produced last week.

NIST asked about this directly. Question 2(b) of the RFI asks "to what degree, if any, could the effectiveness of technical controls, processes, and other practices vary with changes to model capability, agent scaffold software, tool use, deployment method… use in multi-agent systems, and otherwise?" Question 2(d) asks what is involved in "patching or updating AI agent systems throughout the lifecycle, as distinct from those affecting both traditional software systems and non-agentic AI." Both are posed to the field as open questions.

Why is a schedule the wrong trigger?

Here is the structural failure, and it is an architecture problem rather than a discipline problem.

Automated eval suites run on a clock that belongs to the software delivery process: every commit, every nightly build, every release candidate. Anthropic's guidance places them precisely there, as a pre-launch and CI/CD control "running on each agent change and model upgrade." That is the correct recommendation. The difficulty is that CI reliably sees code, and most of those six axes can move without any code being written.

A prompt edited in a dashboard produces no commit. Nor does a tool whose upstream definition is rewritten by its vendor, a permission widened in an admin console to unblock somebody on a Friday, a provider silently routing to a new model snapshot, or an orchestrator reconfigured to spawn a fourth sub-agent. The delivery pipeline sees code. The configuration moves elsewhere, and the suite that would have caught the regression sits idle, because nothing told it anything had happened.

The schedule and the invalidating event therefore run on two different clocks with nothing correlating them, which produces a class of failure that is invisible by construction: the suite is green, the last run was Tuesday, and the system it certified stopped existing on Wednesday afternoon. Anthropic's own comparison table lists this as a standing cost of automated evals — they require ongoing maintenance as the product and the model evolve, or they drift, and they "can create false confidence" once they stop matching real usage. The limitation is published by the people recommending the practice.

Aggregate scoring compounds it. As we have written before, a suite average can hold steady while individual cases rot beneath it, which means a drifting configuration can degrade real behaviour without moving the number anyone watches.

What does change-triggered re-evaluation require?

If the trigger cannot be the calendar, it has to be the change itself. That demands something most eval infrastructure was never asked to provide: each axis needs an observable identity that can be compared across runs.

A model needs a pinned version string, not a provider alias. A prompt needs a content hash. A tool needs a fingerprint of its definition, so that a rewritten upstream description is a detectable event rather than a silent substitution. A permission set needs to be enumerable and diffable. A memory scope needs a stateable boundary. A topology needs recording as the shape it took at execution time, parent and child, rather than the diagram someone drew in onboarding.

Once those identities exist, re-evaluation becomes a function of a diff rather than a cron entry, and a test result gains something it usually lacks: a stated validity condition. This suite passed against this configuration fingerprint; present a different one and the result is stale by definition, and the system knows it.

The other half of CSA's lesson is the environment. Re-running a suite is worth little if the environment has drifted from production, because a green result then certifies a configuration nobody is operating. Production-realistic environments and change-triggered re-evaluation are the same requirement seen from two angles: the test has to describe the deployed system, and it has to re-run whenever the deployed system stops being the thing described. Output-level evals alone reach neither condition — a point worth reading alongside why output evals miss the failures that reach production.

How Waxell handles this

Waxell's Testing capability addresses the fidelity half of the problem. It is a pre-production validation environment that, in its own description, "runs the same policies, budgets, and execution logic as production in an isolated sandbox that cannot mutate production state." Tests point at the real governance plane rather than a parallel copy of it — the page is explicit that there are "no separate test definitions to maintain or keep in sync." That removes the commonest source of environment drift: a fixture that quietly stops resembling the thing it stands in for. Tests run through production's orchestration paths and are designed to exercise boundary conditions, interruptions and failure modes rather than happy paths.

The page is equally explicit about what a passing test means. Every test execution produces a persistent, inspectable result in Waxell Observe — described as "evidence of what was validated, not an assurance that it was." That is the right epistemics for everything above: a result is a dated record of a configuration, and should be read as one.

Observe itself captures every LLM call, tool invocation and agent decision, linking them into full execution trees with parent-child span relationships, OpenTelemetry-native. That gives the topology axis an observable record — the shape a run actually took, rather than the shape it was meant to take. Policies are organised into 50+ categories and evaluated in real time — before execution, between steps and after completion — so the governance behaviour a test exercises is the behaviour that runs in production.

None of this decides when to re-test. It does mean the two hard prerequisites — an environment that matches production, and a durable record of each run's configuration — are infrastructure rather than homework.

FAQ

What is agent testing?

Agent testing is the practice of validating an AI agent system before and during production use, covering not just model outputs but the tools, permissions, environment, orchestration logic and grading harness that surround the model. Because those components can change independently, an agent test result describes one configuration at one moment rather than a permanent property of the agent.

Why do passing agent evals stop being accurate?

Because the system changes on a different schedule than the suite runs. CSA's synthesis of NIST agent-security RFI responses names six axes along which agentic systems evolve — models, prompts, tools, permissions, memory and topology — and recommends that assurance be continuously refreshed rather than treated as a one-time gate. A prompt edit, a widened permission or a rewritten upstream tool definition can each invalidate a result without producing a commit that would re-trigger CI.

What is a production-realistic test environment?

It is an environment that preserves the security-relevant behaviour of deployment — the same policies, permissions, orchestration paths and tool access — without exposing real systems to risk. The phrase comes from CSA's fifth lesson, and the point of it is that an environment which has drifted from production produces results that certify a configuration nobody is running.

Should agent evals run on a schedule or on change?

Both, and the second is usually the one missing. Scheduled runs catch slow drift in model behaviour and grading noise; change-triggered runs catch the discrete events that invalidate a result outright. Anthropic's eval guidance recommends running automated evals on each agent change and model upgrade. The practical obstacle is that several change axes never pass through a code repository, so the trigger has to be built on configuration identity rather than commits.

Can a low eval score mean the eval is broken rather than the agent?

Frequently. Anthropic reports that Claude Opus 4.5 initially scored 42% on CORE-Bench because of rigid grading, ambiguous task specifications and irreproducible stochastic tasks, and reached 95% once those were fixed and the scaffold was loosened. Its guidance also notes that a 0% pass rate across many trials with a frontier model usually signals a broken task rather than an incapable agent. Reading transcripts, rather than trusting the aggregate, is how the difference gets caught.

What should a team do first?

Write down what each of the six axes currently is, and where a change to it would be recorded. The axes that turn out to have no owner and no change log are the ones worth fixing first — that inventory is usually more actionable than another eval score.

Sources

See what your agents are actually doing before you decide what to test — book a 30-minute demo of Waxell Observe: every LLM call, tool invocation and agent decision, captured from two lines of code.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.

Waxell

Waxell provides observability and governance for AI agents in production. Bring your own framework.

Compliance — NIST AI RMF · EU AI Act · SOC 2 compliant · HIPAA (in progress)

Governed continuously in Vanta.

SOC 2 Compliant badge

© 2026 Waxell. All rights reserved.

Patent Pending.