Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Analysis7 min readSample

What the widening benchmark-deployment gap actually tells us

Agent benchmarks are saturating while production success rates lag. The gap is not measurement error — it is the measurement working correctly on the wrong thing.

The Agitology DeskEditorial · Agitology

Systems that score highly on agentic evaluations succeed materially less often in production, and the gap is widening rather than closing. It is tempting to read this as benchmarks being bad. That is not quite it.

Benchmarks specify away the hard part

To be a benchmark, a task must be specified: a clear goal, a defined environment, an unambiguous success criterion. That specification is not overhead — it is most of the difficulty in real work. Production tasks arrive underspecified, in environments that change while you work in them, with success criteria that are contested.

A benchmark therefore measures execution given specification. Deployment measures execution plus specification. As execution saturates, the residual is entirely the part the benchmark removed, which is exactly why the gap widens as scores rise.

The benchmark did not get worse. It got saturated, and what it was never measuring became the whole of what remains.

Why the Index weights deployment above benchmarks

Scaled deployment is reliability evidence no benchmark provides. Millions of people using a system daily generates a distribution of inputs no evaluation set approximates, and commercial deployment reveals failure rates that laboratory conditions suppress.

The cost of this weighting is a lag. Deployment evidence arrives late, so the Index will under-read a genuine capability jump for a quarter or two. That is an acceptable trade against systematically over-reading benchmark saturation, which is the failure mode that has repeatedly embarrassed timeline forecasts.

The reading for capability claims

  • A benchmark result near saturation carries little information about deployment. Treat it as a ceiling, not an estimate.
  • An unsaturated benchmark still discriminates and is worth reading directly.
  • Deployment evidence at scale outweighs both, and it is the evidence labs are least willing to publish.
  • The absence of a field reliability figure, in any dimension, is information.

This is the same shape as the contamination problem in generalization: the measurement is fine, and the thing it measures has stopped being the thing we care about.

Continue

More analysis

AnalysisSample

How to read the Index without fooling yourself

A single number about AGI progress is either a useful summary or a false precision machine. Which one depends entirely on how you read it.

The Agitology Desk6 min

Every Tuesday

The AGI Brief

Five developments, three papers, one forecast, and what moved the Index.

Every Tuesday. No tracking, no sponsor content in the body. Unsubscribe in one click.