Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Analysis9 min readSample

Verification is the bottleneck, and it is not improving

Generation got dramatically cheaper. Checking the output did not. That asymmetry determines where capability gains actually land.

The Agitology DeskEditorial · Agitology

Coding scores 81 in the Index — the second-highest dimension, and the one with the most visible economic footprint. Survey work this month reports that reviewing machine-written code now consumes more time in a majority of engineering organisations than writing it did. Both things are true, and the tension between them is the most useful fact in the Index right now.

The asymmetry

Generating a plausible artefact and confirming a correct one are different problems with different cost curves. The last four years compressed the first curve by orders of magnitude. The second has barely moved, because verification is bounded by the reviewer, and the reviewer is human.

Where verification is cheap — the answer is obviously right or obviously wrong, a test either passes or does not — capability gains convert to value almost directly. Where verification is expensive, they convert into review labour.

0

Realised gain as verification cost approaches generation cost

The model result holds regardless of how capable the generator becomes.

This generalises past software

Software is where the effect is measured because software has instrumentation. The mechanism is not specific to it. Legal drafting, medical documentation, financial analysis, scientific literature review: each has cheap generation and expensive verification, and each shows the same pattern of enthusiastic adoption followed by a much smaller measured gain than expected.

The Index treats this as counter-evidence rather than as a score reduction, and the distinction is deliberate. The capability is real. Its economic realisation is bounded by something the capability does not address.

Why silent failure makes it worse

A system that fails loudly is cheap to verify — the failure announces itself. Current systems mostly fail quietly, producing fluent and confident wrong output. Long-context retrieval work this year found mid-context failures essentially undetectable without ground truth the user does not have.

This is the property that makes autonomy expensive. You cannot remove the human when the failure mode is one only a human notices. It is why autonomy sits at 53 while coding sits at 81: the gap between them is a verification gap, not a capability gap.

What to watch

  1. Calibration: whether systems become reliably able to signal uncertainty at the point of failure rather than after it.
  2. Interpretability coverage: what fraction of a frontier model's computation can be inspected. Currently a small fraction, and oversight decisions need a large one.
  3. Formal verification in domains that admit it — the only place verification cost has ever fallen by an order of magnitude.

None of these is a capability benchmark, which is roughly the point. The next phase of this is not going to be decided by how much better the models get.

Continue

More analysis

AnalysisSample

How to read the Index without fooling yourself

A single number about AGI progress is either a useful summary or a false precision machine. Which one depends entirely on how you read it.

The Agitology Desk6 min

Every Tuesday

The AGI Brief

Five developments, three papers, one forecast, and what moved the Index.

Every Tuesday. No tracking, no sponsor content in the body. Unsubscribe in one click.