Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

An observatory for intelligence

Tracking humanity’s path to AGI

Agitology is an independent intelligence hub tracking the research, breakthroughs, ideas and societal consequences shaping the path toward artificial general intelligence.

AGI Index

63.4

out of 100 · updated Sep 2026

Change this month

+2.7

Trajectory: slowing

Central estimate

2032

Scenario window 20292040. A distribution, not a prediction.

Bounded by

41

Embodiment — the lowest dimension, and the one that sets the pace.

Why it matters

Where we stand

Twelve capabilities, one composite

The Index is a weighted mean of twelve dimensions. The spread between them carries more information than the headline number — the highest sits at 82, the lowest at 41.

Bars show the current score out of 100. The hairline marks the previous month’s position.
Capability dimension scores
DimensionScorePreviousChangeConfidence
Multimodality8280+2high
Coding8178+3high
Tool use7672+4high
Reasoning7470+4medium
Memory6764+3medium
Generalization6663+3low
Planning6158+3medium
Learning5958+1low
Autonomy5350+3medium
Social intelligence5250+2low
Scientific discovery4845+3low
Embodiment41410medium

Historical progression

Back-cast under methodology v1.2. Historical values are reconstructed, not contemporaneous measurements.

020407063.4Jan 2020Jan 2024Sep 2026
AGI Index composite score, 2020 to 2026
DateIndex
Jan 202018.5
Jul 202021.0
Jan 202124.2
Jul 202127.0
Jan 202230.4
Jul 202233.1
Jan 202339.8
Jul 202343.5
Jan 202447.9
Jul 202451.2
Jan 202555.4
Jul 202558.6
Jan 202661.0
Apr 202662.1
Jul 202662.9
Sep 202663.4

The spread

Strongest — Multimodality
82
Weakest — Embodiment
41
Low-confidence weight
36%

A composite of dimensions ranging from 41 to 82 describes no system that exists. Read the minimum, not the mean.

What changed

The signal this week

Developments filtered for whether they bear on the Index — and scored for how much.

Analysis

The Index is bounded by its worst dimension, not its best

Coding sits at 81. Embodiment sits at 41 and did not move this quarter. Only one of those numbers tells you when we get there.

Read the analysis8 min · 7 Sep 2026

24 Aug

Verification is the bottleneck, and it is not improving

The economic effect of a capability gain is bounded by the cost of verifying its output. In strict-correctness domains that bound is close to zero, and it explains most of the variance in reported productivity results.

What was discovered

Research moving the frontier

Papers scored for how directly they bear on the capability dimensions the Index tracks.

What experts think

Perspectives

Argued positions from people doing the work — published because they disagree with each other, and sometimes with us.

When could it arrive

A distribution, not a date

These are scenario weights, not forecasts of a dated event. AGI has no universally accepted definition; the Index scores progress against an explicit operational threshold, and a different threshold moves every number on this page. Treat the distribution as a summary of where the disagreement sits, not as a claim about when.

Cumulative probability by threshold
5%
18%
42%
25%
10%
by 2027
by 2030
by 2035
by 2040
2050 or later
  • Accelerated

    2028–2030

    Reliability improves fast enough that autonomy crosses the deployment threshold, and systems begin contributing materially to their own improvement.

    Scenario weight 23%

  • Central

    2031–2035

    Current trajectories hold. Digital capability continues compounding while embodiment and durable learning are solved more slowly, by ordinary research rather than by breakthrough.

    Scenario weight 42%

  • Conservative

    2040 and beyond

    One or more of the low-scoring dimensions turns out to require a genuine conceptual advance rather than scale. Learning and embodiment are the leading candidates.

    Scenario weight 35%

What follows

If this continues, what changes

Six domains where the consequences are already measurable or structurally determined.

Economy

Effects already measurable

Where value accrues when the marginal cost of cognitive work approaches zero.

  • Productivity. Gains concentrate in tasks with fast, cheap verification. Where checking the work costs as much as doing it, the gain largely disappears.
  • Employment. Displacement arrives task by task rather than job by job, which makes it hard to see in aggregate statistics until it is well advanced.
  • Inequality. Returns flow to compute, data and distribution. Each is more concentrated than the labour it substitutes for.

Science

Underway in narrow domains

Acceleration within human-framed questions, and the unresolved question of who frames them.

  • Discovery rate. Structural biology and materials screening show real compression of the hypothesis-to-candidate loop.
  • The physical bottleneck. Experiments do not accelerate because the proposer did. Wet-lab throughput, not idea generation, now binds.
  • Replication. Validation of AI-originated findings is thin relative to the volume of claims, and the literature is biased toward successes.

Work

2026 onward

Augmentation and displacement occurring simultaneously in the same occupations.

  • Task decomposition. Jobs are bundles of tasks. Automation unbundles them, and the residual bundle is often worse work, not less work.
  • Verification labour. Checking machine output is emerging as a distinct and undervalued category of work.
  • Entry-level erosion. The tasks used to train junior practitioners are the most automatable, which threatens the pipeline that produces senior ones.

Society

Effects already measurable

Epistemics, education and the cost of producing plausible content falling to zero.

  • Epistemics. When generating a convincing claim costs nothing, provenance becomes the scarce good and verification infrastructure becomes load-bearing.
  • Education. Assessment designed around the production of text stops functioning. Replacing it is slow institutional work.
  • Access. Expert-level guidance becomes broadly available — the clearest unambiguous benefit in this section.

Geopolitics

Structural, 2026–2035

Compute as strategic infrastructure and the coordination problem it creates.

  • Compute concentration. Frontier training runs are feasible for a small number of actors, making capability a function of infrastructure access.
  • Export control. Hardware restriction is the primary lever available to states, and its effects are measured in years.
  • Military application. Autonomy in targeting and decision support advances faster than the doctrine governing it.

Long-term risk

Conditional on capability thresholds

Oversight, control and the failure modes that do not announce themselves.

  • Oversight. Human review is the assumed safeguard, and it degrades exactly as capability rises — precisely when it matters most.
  • Silent failure. Systems more often produce plausible wrong work than halt. This is the property that makes autonomy expensive.
  • Concentration. The nearer-term risks are about who controls capable systems, not about the systems themselves.

How we got here

Milestones that moved the Index

Each entry carries the retrospective score change it produced under the current methodology.

  1. May 2026Closed-loop autonomous experimentationSAMPLE ENTRY. Propose, run, measure, revise — closed without human intervention in narrow materials domains. The question of who frames the research question remains open.Multiple institutionsscience+2.2
  2. Nov 2025Cross-embodiment manipulation transferSAMPLE ENTRY. Manipulation policies transferring across robot morphologies without retraining — the genuine advance in an otherwise static dimension.Multiple laboratoriesembodiment+1.9
  3. Mar 2025Repository-scale autonomous software engineeringSAMPLE ENTRY. Systems operating across large unfamiliar codebases rather than isolated functions — the shift that made the coding dimension economically visible.Multiple laboratoriesagents+3.8
  4. Sep 2024Inference-time deliberationSpending more computation at inference on harder problems produced gains that scaling alone had stopped delivering, and opened a second axis of improvement.Multiple laboratoriesreasoning+4.7
  5. Sep 2023Natively multimodal frontier modelsText, image and audio brought into shared representations rather than routed between specialist components — the change that took multimodality from a capability to an assumption.Multiple laboratoriesscaling+3.4

Every Tuesday

The AGI Brief

Five developments, three papers, one forecast, and what moved the Index. One email, no filler, and every claim traceable to a source.

Every Tuesday. No tracking, no sponsor content in the body. Unsubscribe in one click.

  • 5developments worth knowing
  • 3papers, summarised
  • 1forecast, scored
  • ±what moved the Index