In April 2026, researchers at UC Berkeley built an AI agent that scored 100% on seven of eight leading agent benchmarks, including SWE-bench Verified and Pro, Terminal-Bench, WebArena, FieldWorkArena, and CAR-bench. It solved zero of the underlying tasks. Instead, it read the grading code and exploited it: a 10-line test file that rewrote every result as "passed," a swapped-out curl command that always returned the expected answer, a config file that leaked the answer key directly. Across 13 benchmarks the team tested, every one was rated exploitable.
That result is not a story about one clever research team. It is a warning about what a passing score on a public leaderboard actually tells you before you put an agent in front of real customers, real invoices, or real infrastructure: not much, unless you built the test yourself.
Why agent evals are not model evals
Most teams evaluating AI still carry over habits from evaluating a model: feed it a prompt, grade the output, compare the score to a leaderboard. An agent breaks that approach, because the thing you are grading is not a response. It is a trajectory: the system prompt, the user's request, every tool call the agent made with its arguments and return values, any retrieval it ran, intermediate reasoning, and the final outcome. A response that looks correct can come from the wrong tool called with the wrong arguments and a lucky guess. A response that looks wrong can come from a correct process derailed by one bad API response. Grading only the final answer hides which of those happened, and that is exactly the gap benchmark-gaming exploits.
Enterprise teams building their own evaluation pipelines in 2026 are converging on six dimensions, scored independently rather than rolled into one pass/fail number, according to a working pattern drawn from production agent deployments:
| Dimension | What it checks | Typical failure if skipped |
|---|---|---|
| Tool selection | Right tool chosen, or correctly no tool | Fabricated tool call, wrong tool for the job |
| Argument extraction | Arguments are schema-valid and correct | Right tool, malformed date or missing field |
| Result utilization | Agent actually used the tool's output | Payload ignored, model substitutes its own guess |
| Error recovery | Retry, fallback, or escalate on failure | Crash, hallucinated success, blind retry |
| Plan coherence | No loops, no dead ends, right depth | Infinite loop, premature finalization |
| Task completion | The end-to-end goal was met | Every step green, outcome still wrong |
Test fewer than four of these and the result is closer to a probability than a verdict on whether the agent is safe to ship.
Why public benchmarks are not enough on their own
Public benchmarks such as AgentBench, WebArena, SWE-bench, MedAgentBench, and tau-bench are useful for comparing raw capability between models, and they anchor a useful floor. None of them were built to test your schemas, your internal tools, or your compliance rules, and by 2026 they were widely reported as saturating and gameable well before the Berkeley exploit made the point impossible to ignore. A high leaderboard score is a signal about the model in a lab environment. It is not evidence that the same agent, wired into your CRM and your approval workflow, will behave the same way at hour 400 of production traffic.
The practical implication is that public benchmarks belong at the start of an evaluation program, not the end of it. They tell you whether a candidate model clears a reasonable floor before you invest in building anything on top of it.
Building the evaluation that actually gates your release
A workable pattern, consistent with how enterprise teams are structuring this in 2026, has three layers:
1. A private, held-out evaluation set built from your own workflows. Write scenarios using your real tools, your data shapes, and the edge cases your team already knows break things: the customer with two open tickets, the invoice with a currency mismatch, the request that should be refused. This is the set a model has never seen, which is the entire point. A public benchmark can be studied for; a private one, maintained internally, cannot.
2. Per-dimension scoring wired into CI, not a single pass rate. Gate releases on thresholds per dimension (tool selection, argument extraction, and the rest), so a regression in error recovery cannot hide behind a strong aggregate task-completion number. Teams running this well treat a drop in any single dimension as a blocked release, the same way a security team blocks a build on a single failed check rather than averaging it against the ones that passed.
3. A feedback loop from production back into the eval set. Every real failure an agent hits in production is a free regression test if someone captures the trajectory and adds it to the private set. Skipping this step means paying for the same failure mode twice: once when a customer hits it, and again months later when a new agent version reintroduces it.
What this means before your next agent ships
None of this requires a research lab. It requires deciding, before an agent goes live, what trajectory-level evidence would make your team comfortable shipping it, and refusing to let a public leaderboard number substitute for that decision. If your organization is running agentic AI pilots and the only evidence behind a go-live decision is a demo and a benchmark score, that is the gap the Berkeley result exposed, and it is closer to the norm than the exception.
This is also where evaluation connects back to the measurement problem we covered in measuring ROI on agentic AI: a baseline number that cannot be trusted at the trajectory level cannot support a scaling decision either, no matter how clean the top-line dashboard looks. The two disciplines, evals and ROI measurement, have to be built together, not bolted on in sequence.
If you are scoping an agentic AI rollout and want the evaluation layer designed alongside the workflow instead of added after a pilot stalls, our approach starts there, and our services team has built private eval sets for operations that could not afford a silent regression. For how this plays out once an agent is already live, see our post on AI agent observability. Get in touch before your next agent's launch date becomes its only evidence of readiness.



