Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

Methodology v1.2

How the AGI Index is built

AGI cannot be objectively measured. There is no agreed definition, no unit, and no instrument. What follows is not a claim to have solved that — it is a transparent framework that produces a number, stated clearly enough that you can disagree with it precisely.

The threshold being scored against

Every number on this site is conditional on one operational definition. We score progress toward a system that can, without task-specific engineering:

  • Perform the large majority of economically valuable cognitive work at or above competent human level;
  • Transfer competence to problems it was not built for, including physical ones;
  • Operate over long horizons without supervision, and recognise when it is failing;
  • Acquire durable new capability after deployment.

This is a demanding threshold, deliberately. A weaker one — say, matching human performance on a fixed benchmark suite — would already be near-satisfied, and would make the Index useless. A stronger one, requiring recursive self-improvement or consciousness, is unfalsifiable. Change the threshold and every number here moves. That is a property of the subject, not a flaw in the framework.

Composition

Twelve dimensions, each scored 0–100 and each carrying a fixed weight. The composite is the weighted mean — computed from the dimension scores at build time, never stored. Editing any dimension moves the headline number, so the composite cannot drift away from the components that are supposed to justify it.

Weights reflect how load-bearing each capability is for the threshold above, not how interesting or well-measured it is. Embodiment carries 7% despite being the hardest to evidence, because a system that cannot act in the physical world does not meet the definition.

What counts as evidence

Sources are weighted in this order, strongest first:

  1. Deployment at scale. Millions of users generate an input distribution no evaluation approximates, and commercial deployment surfaces failure rates that laboratory conditions suppress.
  2. Independent replication. Especially where the replicating group has no stake in the result.
  3. Contamination-resistant evaluation. Held-out items constructed after the relevant training cut-off.
  4. Published benchmarks. Useful while unsaturated; nearly uninformative once approaching ceiling.
  5. Demonstrations. Weighted close to zero. A demonstration is a selected best run.

This ordering has a cost: deployment evidence arrives late, so the Index will under-read a genuine capability jump for a quarter or two. We accept that lag rather than over-read benchmark saturation, which is the error that has repeatedly embarrassed timeline forecasting.

Confidence, separately from score

Each dimension carries a confidence rating that is independent of its score. A well-evidenced 74 and a barely-evidenced 59 are different claims, and the composite cannot show that difference — which is why the dimension table does.

Currently 36% of the composite weight sits on low-confidence dimensions. That is where the real uncertainty in the headline number lives.

Counter-evidence is published

Every dimension page lists what holds the score down alongside what raises it, at equal prominence. A framework that only records supporting evidence is not measuring anything, and scores have moved downward twice this year on the strength of it.

Back-casting, and its limits

Historical values are reconstructed under the current methodology version. Nobody was running this Index in 2017, so the history is a back-cast, not a record. Back-casts are systematically flattering to the framework producing them — the shape of the curve is partly a property of the method. Treat the trend as indicative and the current value as the claim.

Forecasts

The distribution on the Index page is a set of scenario weights, not a prediction of a dated event. Scenarios are defined by what would have to be true rather than by when, and each lists its own preconditions so you can score it yourself as evidence arrives.

Revisions

Score changes are published with their reason. Methodology changes increment the version and trigger a full back-cast, which is disclosed rather than applied silently. A number that only ever moves up is a marketing instrument.

What this framework cannot do

  • It cannot tell you the date. It summarises where disagreement sits.
  • It cannot escape its own threshold. Every figure is conditional on the definition above.
  • It cannot fully rule out contamination in the evidence it reads from.
  • It cannot weight dimensions objectively. The weights are argued, not derived.