AGI Index dimension · weight 9%
Autonomy
Useful operation without a human in the loop.
How long a system can run unsupervised on a real task before human intervention is required to keep it correct, safe or on-budget. Measured in wall-clock time and in consequential decisions taken without review.
The argument
Evidence and counter-evidence
Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.
Raises the score
- Unsupervised operating windows have lengthened from minutes to hours on well-specified engineering work.
- Self-monitoring — recognising that a run has gone off the rails and stopping — is beginning to appear.
Holds it down
- Deployment practice remains overwhelmingly supervised, which is itself evidence about trust.
- Failure modes are silent: systems more often produce plausible wrong work than halt, which raises the cost of removing the human.
- Autonomy is measured almost entirely in domains with cheap, fast feedback. It is largely untested where errors are expensive.
What the score reads from
Unsupervised operating window
Intervention rate
Related coverage
Everything bearing on autonomy
News
- Unsupervised operating windows reach a full working day on constrained tasksAnthropic · 4 Sep
- Closed-loop materials laboratory completes a month without human interventionUniversity consortium · 30 Aug
- Oversight requirements tighten for autonomous deployment in regulated sectorsEuropean Commission · 25 Aug
- Verification, not generation, emerges as the primary bottleneck in software teamsIndustry survey · 22 Aug
- Interpretability tooling scales to frontier-size modelsAnthropic · 9 Aug
Research
- The Divergence Between Agent Benchmarks and Deployment OutcomesPreprint · 11 Jul
- Closed-Loop Autonomous Experimentation in Materials DiscoveryNature Machine Intelligence · 15 Jun
- Feature-Level Interpretability at Frontier ScalePreprint · 22 May
- Unsupervised Operating Windows in Production Agent DeploymentsPreprint · 3 Apr
Argument
- Verification is the bottleneck, and it is not improvingAnalysis · 24 Aug
- What the widening benchmark-deployment gap actually tells usAnalysis · 11 Aug
- Scientific discovery: everything except the part that mattersAnalysis · 16 Jul
- The demo-to-deployment gap is not a detail, it is the problemPerspective · 19 Aug
Context
Where this sits against the rest
- Multimodality82
- Coding81
- Tool use76
- Reasoning74
- Memory67
- Generalization66
- Planning61
- Learning59
- Social intelligence52
- Scientific discovery48
- Embodiment41
How these scores are producedMethodology v1.2