Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Perspective9 min readSample

We are measuring recall and calling it reasoning

Until evaluations are built to survive models that may have seen them, capability claims are unfalsifiable — and the field has not paid the cost of fixing this.

The position

Most published capability gains cannot currently be distinguished from contamination, and the field has chosen not to find out.

Dr. Renata OkonkwoResearch scientist, evaluation · Independent

I want to state the problem as plainly as I can. We do not know whether frontier systems generalize, because we have not built evaluations capable of telling us. Every widely cited transfer benchmark predates the models it is used to evaluate, and every one of those models was trained on a corpus we cannot fully inspect.

This is not a hypothetical concern. When my colleagues and I rebuilt six standard transfer benchmarks using items authored after the relevant training cut-offs, matched for difficulty, measured transfer fell substantially on all six. The drops were largest on the benchmarks most frequently cited in capability announcements.

Why this is not fixable with better filtering

The usual response is decontamination: filter the test set out of the training data. This fails for two reasons.

First, exact-match filtering does not catch paraphrase, translation, or discussion of a problem, and discussion is everywhere. A benchmark item that has been blogged about, argued over on forums and worked through in tutorials is thoroughly present in the corpus without appearing verbatim once.

Second, and more fundamentally, decontamination is performed by the organisation that also benefits from the resulting number. I do not think anyone is cheating. I think the incentive gradient is obvious and nobody should be asked to audit themselves on a question this consequential.

An evaluation that cannot fail is not an evaluation. It is a press release with error bars.

What would actually work

  1. Items authored after training cut-offs, by people with no stake in the result, held privately and used once.
  2. Difficulty matched against a public set so the comparison means something.
  3. Published construction criteria, so the evaluation can itself be criticised.
  4. Single use. A held-out set that is reused becomes a training set the moment it is reported on.

This is expensive. Item authoring does not scale, single-use sets are consumed by definition, and it is unrewarding work — nobody is promoted for building a benchmark that makes the numbers go down.

What I am not saying

I am not saying frontier systems do not reason. I use them daily and I think something real is happening. I am saying that the evidence we currently produce cannot distinguish between the hypothesis that they reason and the hypothesis that they retrieve very well, and that a field which cannot distinguish its central hypothesis from its most obvious null is not yet doing measurement.

The Index rates generalization at low confidence. That is the correct rating, and it should stay there until somebody builds the evaluation that would let it change.

Continue

Other perspectives