AGI Index dimension · weight 7%
Multimodality
Fluent operation across text, image, audio and video.
Whether a single system perceives and produces across modalities with shared representations, rather than routing between specialised components. The most mature dimension in the Index.
The argument
Evidence and counter-evidence
Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.
Raises the score
- Unified architectures handle text, image, audio and video without task-specific heads, and reason across them in a single pass.
- Real-time interactive audio and video is deployed at consumer scale, which is a strong reliability signal.
Holds it down
- Fine-grained spatial and temporal grounding remains weak — counting, ordering and precise localisation still fail.
- Cross-modal reasoning lags cross-modal perception: systems see the scene and misjudge what it implies.
What the score reads from
Cross-modal understanding
Spatial and temporal grounding
Multimodality over time
| Date | Index |
|---|---|
| Jan 2020 | 12.0 |
| Jul 2020 | 15.0 |
| Jan 2021 | 21.0 |
| Jul 2021 | 27.0 |
| Jan 2022 | 34.0 |
| Jul 2022 | 41.0 |
| Jan 2023 | 48.0 |
| Jul 2023 | 55.0 |
| Jan 2024 | 63.0 |
| Jul 2024 | 69.0 |
| Jan 2025 | 74.0 |
| Jul 2025 | 78.0 |
| Jan 2026 | 80.0 |
| Apr 2026 | 80.0 |
| Jul 2026 | 81.0 |
| Sep 2026 | 82.0 |
Related coverage
Everything bearing on multimodality
Context
Where this sits against the rest
- Coding81
- Tool use76
- Reasoning74
- Memory67
- Generalization66
- Planning61
- Learning59
- Autonomy53
- Social intelligence52
- Scientific discovery48
- Embodiment41
How these scores are producedMethodology v1.2