Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Analysis8 min readSample

The Index is bounded by its worst dimension, not its best

Coding sits at 81. Embodiment sits at 41 and did not move this quarter. Only one of those numbers tells you when we get there.

The Agitology DeskEditorial · Agitology

The Index reads 63.4. It is a weighted mean, and like every weighted mean it hides its own distribution. The twelve dimensions underneath it range from 41 to 82 — a spread wide enough that the composite describes no actual system anyone has built.

That spread is the finding. A threshold defined by generality is not reached by averaging over capabilities; it is reached when the weakest necessary capability clears the bar. On that reading, the number to watch is not 63.4. It is 41.

41

Embodiment score

Unchanged for two consecutive quarters — the only dimension in the Index that is flat.

The digital dimensions are compounding

Coding, multimodality and tool use — 81, 82 and 76 — all share a property: abundant training data, cheap evaluation, fast feedback. They are the dimensions where the field can iterate quickly, and they have improved accordingly. Reasoning at 74 is close behind and moving fastest of the four.

It is tempting to read that cluster as the leading edge and assume the rest follows. The history of the field offers some support: capabilities that looked distinct have repeatedly turned out to share machinery, and progress in one has arrived in another without anyone targeting it.

The physical dimensions are not

Embodiment does not share that property, and the reason is structural rather than incidental. Physical interaction data cannot be scraped. Every hour of it costs an hour of a robot moving in the world, plus the hardware to move and the space to move in. No amount of capital compresses that below the speed of physical time.

Cross-embodiment transfer is the genuine advance here, and it attacks precisely this constraint by letting data collected on one platform serve others. It is the reason the dimension has a positive outlook despite a flat score. But laboratory transfer is not field reliability, and field reliability numbers are conspicuously unpublished across the industry.

Learning is the quieter problem

Embodiment at least has a visible research programme with a legible constraint. Learning, at 59, is stranger. The score looks middling, which understates how little of it reflects the thing the dimension is supposed to measure.

Almost all of the improvement in learning has come from building better scaffolding around static weights: retrieval, context, memory systems. These work. They are also, straightforwardly, not learning. Nothing acquired in deployment persists into the model. Replication work this year found that continual-learning methods effective at research scale stop working at frontier scale, which suggests the problem is not close to solved and that the field has been engineering around it rather than at it.

A capability you have engineered around is not a capability you have acquired. The distinction matters most exactly when you try to remove the scaffolding.

What this does to the timeline

It is the main reason the central scenario sits at 2031–2035 rather than earlier. The accelerated scenario requires autonomy above 75 and durable learning, and both currently depend on problems where progress is engineering-shaped rather than breakthrough-shaped. Engineering is reliable and it is slow.

  • If embodiment and learning are scale problems, the accelerated scenario is live and the Index is under-reading.
  • If either requires a conceptual advance, the conservative scenario dominates and the digital dimensions will keep rising against a ceiling that does not move.
  • The evidence does not currently distinguish between these, which is why the forecast confidence is medium and not high.

Read the composite if you want one number. Read the minimum if you want the answer.

Continue

More analysis

AnalysisSample

How to read the Index without fooling yourself

A single number about AGI progress is either a useful summary or a false precision machine. Which one depends entirely on how you read it.

The Agitology Desk6 min

Every Tuesday

The AGI Brief

Five developments, three papers, one forecast, and what moved the Index.

Every Tuesday. No tracking, no sponsor content in the body. Unsubscribe in one click.