AGI Index dimension · weight 7%
Coding
Writing, reading and repairing software.
Software engineering as practised rather than as benchmarked: navigating an unfamiliar codebase, making a change that does not break adjacent behaviour, and explaining why. The most economically visible dimension, which makes it the easiest to over-read.
The argument
Evidence and counter-evidence
Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.
Raises the score
- Resolution rates on real repository issues have risen steeply and are now a routine part of professional workflow.
- Systems operate across large unfamiliar codebases, not just isolated functions — the shift that made the capability economically real.
Holds it down
- Strong performance concentrates in well-represented languages and idioms; it thins out quickly at the edges.
- Architectural judgement lags implementation skill by a wide margin. Systems write the code well and choose the wrong thing to build.
- Verification, not generation, is now the bottleneck — and it is the part that has moved least.
What the score reads from
Repository-level issue resolution
Architectural decision quality
Coding over time
| Date | Index |
|---|---|
| Jan 2020 | 20.0 |
| Jul 2020 | 26.0 |
| Jan 2021 | 33.0 |
| Jul 2021 | 38.0 |
| Jan 2022 | 43.0 |
| Jul 2022 | 47.0 |
| Jan 2023 | 56.0 |
| Jul 2023 | 61.0 |
| Jan 2024 | 66.0 |
| Jul 2024 | 70.0 |
| Jan 2025 | 73.0 |
| Jul 2025 | 76.0 |
| Jan 2026 | 78.0 |
| Apr 2026 | 78.0 |
| Jul 2026 | 80.0 |
| Sep 2026 | 81.0 |
Related coverage
Everything bearing on coding
Context
Where this sits against the rest
- Multimodality82
- Tool use76
- Reasoning74
- Memory67
- Generalization66
- Planning61
- Learning59
- Autonomy53
- Social intelligence52
- Scientific discovery48
- Embodiment41
How these scores are producedMethodology v1.2