AGI Index dimension · weight 12%
Generalization
Transfer of competence to genuinely unfamiliar problems.
Whether capability learned in one domain transfers to a domain the system was not trained for, without task-specific scaffolding. The dimension most directly implicated by the word "general" in AGI, and the hardest to measure because contamination is difficult to rule out.
The argument
Evidence and counter-evidence
Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.
Raises the score
- Cross-domain transfer is observable: capability acquired in code appears to improve structured reasoning in unrelated symbolic domains.
- Few-shot adaptation to novel task formats has improved substantially without fine-tuning.
Holds it down
- Training-set contamination cannot be excluded for most public evaluations, so an unknown share of apparent transfer is recall.
- Performance on deliberately novel abstraction tasks remains far below human baseline despite enormous gains elsewhere.
- Transfer is asymmetric — it flows readily between text-adjacent domains and poorly into anything requiring physical intuition.
What the score reads from
Abstraction and reasoning corpora
Held-out task formats
Related coverage
Everything bearing on generalization
News
Research
Context
Where this sits against the rest
- Multimodality82
- Coding81
- Tool use76
- Reasoning74
- Memory67
- Planning61
- Learning59
- Autonomy53
- Social intelligence52
- Scientific discovery48
- Embodiment41