Deliberation budgets scale past the point of diminishing returns
A frontier laboratory reports that allocating substantially more inference compute to hard problems continues to yield gains well beyond where the curve was expected to flatten.
PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means
An observatory for intelligence
Agitology is an independent intelligence hub tracking the research, breakthroughs, ideas and societal consequences shaping the path toward artificial general intelligence.
Where we stand
The Index is a weighted mean of twelve dimensions. The spread between them carries more information than the headline number — the highest sits at 82, the lowest at 41.
| Dimension | Score | Previous | Change | Confidence |
|---|---|---|---|---|
| Multimodality | 82 | 80 | +2 | high |
| Coding | 81 | 78 | +3 | high |
| Tool use | 76 | 72 | +4 | high |
| Reasoning | 74 | 70 | +4 | medium |
| Memory | 67 | 64 | +3 | medium |
| Generalization | 66 | 63 | +3 | low |
| Planning | 61 | 58 | +3 | medium |
| Learning | 59 | 58 | +1 | low |
| Autonomy | 53 | 50 | +3 | medium |
| Social intelligence | 52 | 50 | +2 | low |
| Scientific discovery | 48 | 45 | +3 | low |
| Embodiment | 41 | 41 | 0 | medium |
Historical progression
Back-cast under methodology v1.2. Historical values are reconstructed, not contemporaneous measurements.
| Date | Index |
|---|---|
| Jan 2020 | 18.5 |
| Jul 2020 | 21.0 |
| Jan 2021 | 24.2 |
| Jul 2021 | 27.0 |
| Jan 2022 | 30.4 |
| Jul 2022 | 33.1 |
| Jan 2023 | 39.8 |
| Jul 2023 | 43.5 |
| Jan 2024 | 47.9 |
| Jul 2024 | 51.2 |
| Jan 2025 | 55.4 |
| Jul 2025 | 58.6 |
| Jan 2026 | 61.0 |
| Apr 2026 | 62.1 |
| Jul 2026 | 62.9 |
| Sep 2026 | 63.4 |
The spread
What changed
Developments filtered for whether they bear on the Index — and scored for how much.
A frontier laboratory reports that allocating substantially more inference compute to hard problems continues to yield gains well beyond where the curve was expected to flatten.
Analysis
Coding sits at 81. Embodiment sits at 41 and did not move this quarter. Only one of those numbers tells you when we get there.
What was discovered
Papers scored for how directly they bear on the capability dimensions the Index tracks.
Gains from additional inference compute persist across two further orders of magnitude on long-chain problems.
Reported transfer falls materially when evaluation items are rebuilt to exclude training overlap.
A single policy transfers across structurally different robot platforms with limited degradation.
What experts think
Argued positions from people doing the work — published because they disagree with each other, and sometimes with us.
“Most published capability gains cannot currently be distinguished from contamination, and the field has chosen not to find out.”
Until evaluations are built to survive models that may have seen them, capability claims are unfalsifiable — and the field has not paid the cost of fixing this.
“Reliability, not capability, is the binding constraint on autonomy — and it is not on the trajectory the capability curves suggest.”
Four years of shipping agents into environments where mistakes cost money. What actually breaks is never what the benchmarks measure.
“Labour markets have absorbed technological shocks before, but always over a generation. The adjustment speed, not the endpoint, is what is different here.”
The debate is about the eventual equilibrium. What determines whether this is survivable is the speed of the transition, and almost nobody is modelling it.
When could it arrive
These are scenario weights, not forecasts of a dated event. AGI has no universally accepted definition; the Index scores progress against an explicit operational threshold, and a different threshold moves every number on this page. Treat the distribution as a summary of where the disagreement sits, not as a claim about when.
Reliability improves fast enough that autonomy crosses the deployment threshold, and systems begin contributing materially to their own improvement.
Scenario weight 23%
Current trajectories hold. Digital capability continues compounding while embodiment and durable learning are solved more slowly, by ordinary research rather than by breakthrough.
Scenario weight 42%
One or more of the low-scoring dimensions turns out to require a genuine conceptual advance rather than scale. Learning and embodiment are the leading candidates.
Scenario weight 35%
What follows
Six domains where the consequences are already measurable or structurally determined.
Where value accrues when the marginal cost of cognitive work approaches zero.
Acceleration within human-framed questions, and the unresolved question of who frames them.
Augmentation and displacement occurring simultaneously in the same occupations.
Epistemics, education and the cost of producing plausible content falling to zero.
Compute as strategic infrastructure and the coordination problem it creates.
Oversight, control and the failure modes that do not announce themselves.
How we got here
Each entry carries the retrospective score change it produced under the current methodology.