Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

Methodology v1.2 · updated Sep 2026

How close are we to artificial general intelligence?

The Index is a weighted composite of twelve capability dimensions, scored against an explicit operational threshold. It is an analytical estimate under a published framework — not a measurement, because no measurement of this exists.

63.4

Composite score out of 100

+2.7 since Aug 2026

Trajectory

slowing

Second derivative of the last four snapshots.

Forecast confidence

medium

How much any single score should move your beliefs.

Central estimate

2032

Scenario window 2029–2040.

Low-confidence weight

36%

Share of the composite resting on thin evidence.

Progression

The Index over time

Reconstructed under the current methodology version. Because the framework changed, historical values are back-cast rather than contemporaneous — a limitation we would rather state than hide.

020407063.4Jan 2020Jan 2024Sep 2026
AGI Index composite score, 2020 to 2026
DateIndex
Jan 202018.5
Jul 202021.0
Jan 202124.2
Jul 202127.0
Jan 202230.4
Jul 202233.1
Jan 202339.8
Jul 202343.5
Jan 202447.9
Jul 202451.2
Jan 202555.4
Jul 202558.6
Jan 202661.0
Apr 202662.1
Jul 202662.9
Sep 202663.4

Component scores

Twelve dimensions

Sorted by score. The hairline on each bar marks last month's position. Every dimension links to its evidence, its counter-evidence and what would change it.

Bars show the current score out of 100. The hairline marks the previous month’s position.
Capability dimension scores
DimensionScorePreviousChangeConfidence
Multimodality8280+2high
Coding8178+3high
Tool use7672+4high
Reasoning7470+4medium
Memory6764+3medium
Generalization6663+3low
Planning6158+3medium
Learning5958+1low
Autonomy5350+3medium
Social intelligence5250+2low
Scientific discovery4845+3low
Embodiment41410medium

Fluent operation across text, image, audio and video.

high confidenceweight 7%

Writing, reading and repairing software.

high confidenceweight 7%

Operating external systems to extend capability.

high confidenceweight 6%

Multi-step inference that holds together over long chains.

medium confidenceweight 13%

Retention and retrieval across long interactions.

medium confidenceweight 6%

Transfer of competence to genuinely unfamiliar problems.

low confidenceweight 12%

Goal decomposition and recovery when a plan fails.

medium confidenceweight 9%

Acquiring new capability after training ends.

low confidenceweight 9%

Useful operation without a human in the loop.

medium confidenceweight 9%

Generating findings that survive independent replication.

low confidenceweight 9%

Acting competently in unstructured physical space.

medium confidenceweight 7%

Forecast

When could AGI arrive?

Scenario weights, not a prediction of a dated event. A different operational threshold moves every number on this page.

5%
18%
42%
25%
10%
by 2027Discontinuous
by 2030Accelerated
by 2035Central
by 2040Extended
2050 or laterConservative
Scenario probability distribution
ThresholdScenarioProbability
by 2027Discontinuous5%
by 2030Accelerated18%
by 2035Central42%
by 2040Extended25%
2050 or laterConservative10%
  • Accelerated

    2028–2030 · 23%

    Reliability improves fast enough that autonomy crosses the deployment threshold, and systems begin contributing materially to their own improvement.

    Requires

    • Autonomy above 75 — unsupervised operation measured in days, not hours
    • Continual learning that persists into the model, not the scaffolding
    • Verification improving at least as fast as generation
  • Central

    2031–2035 · 42%

    Current trajectories hold. Digital capability continues compounding while embodiment and durable learning are solved more slowly, by ordinary research rather than by breakthrough.

    Requires

    • No sustained plateau in reasoning or generalization
    • Compute and energy supply keeping pace with demand
    • Embodiment closing enough to stop bounding the composite
  • Conservative

    2040 and beyond · 35%

    One or more of the low-scoring dimensions turns out to require a genuine conceptual advance rather than scale. Learning and embodiment are the leading candidates.

    Requires

    • Generalization plateaus as contamination-corrected evaluations mature
    • Continual learning resists engineering solutions
    • Physical-interaction data remains the binding constraint on embodiment

Retrospective

Milestones and their Index impact

Scored retrospectively under the current framework. These are not contemporaneous measurements — nobody was running this Index in 2017.

  1. May 2026Closed-loop autonomous experimentationSAMPLE ENTRY. Propose, run, measure, revise — closed without human intervention in narrow materials domains. The question of who frames the research question remains open.Multiple institutionsscience+2.2
  2. Nov 2025Cross-embodiment manipulation transferSAMPLE ENTRY. Manipulation policies transferring across robot morphologies without retraining — the genuine advance in an otherwise static dimension.Multiple laboratoriesembodiment+1.9
  3. Mar 2025Repository-scale autonomous software engineeringSAMPLE ENTRY. Systems operating across large unfamiliar codebases rather than isolated functions — the shift that made the coding dimension economically visible.Multiple laboratoriesagents+3.8
  4. Sep 2024Inference-time deliberationSpending more computation at inference on harder problems produced gains that scaling alone had stopped delivering, and opened a second axis of improvement.Multiple laboratoriesreasoning+4.7
  5. Sep 2023Natively multimodal frontier modelsText, image and audio brought into shared representations rather than routed between specialist components — the change that took multimodality from a capability to an assumption.Multiple laboratoriesscaling+3.4
  6. Nov 2022Conversational deployment at consumer scaleMoved frontier capability from research artefact to daily instrument for hundreds of millions of people. The Index treats scaled deployment as reliability evidence that no benchmark provides.OpenAIscaling+5.1
  7. Nov 2020Protein structure prediction at experimental accuracyThe first case of a machine learning system resolving a long-standing open problem in the natural sciences, and still the strongest single piece of evidence in the scientific-discovery dimension.Google DeepMindscience+2.9
  8. May 2020Large-scale few-shot language modellingDemonstrated that capability could emerge from scale without task-specific training, reframing the research agenda around scaling rather than architecture.OpenAIscaling+3.6
  9. Jun 2017The transformer architecture is publishedAttention replaced recurrence, making training parallelisable and turning scale into a lever the field could actually pull. Nearly every system the Index tracks descends from this paper.Googlefoundations+4.2