Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

AGI Index dimension · weight 7%

Multimodality

Fluent operation across text, image, audio and video.

Whether a single system perceives and produces across modalities with shared representations, rather than routing between specialised components. The most mature dimension in the Index.

The argument

Evidence and counter-evidence

Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.

Raises the score

  • Unified architectures handle text, image, audio and video without task-specific heads, and reason across them in a single pass.
  • Real-time interactive audio and video is deployed at consumer scale, which is a strong reliability signal.

Holds it down

  • Fine-grained spatial and temporal grounding remains weak — counting, ordering and precise localisation still fail.
  • Cross-modal reasoning lags cross-modal perception: systems see the scene and misjudge what it implies.

What the score reads from

  • Cross-modal understanding

    Approaching human baseline on standard sets.

  • Spatial and temporal grounding

    The clearest remaining gap in an otherwise mature dimension.

Multimodality over time

0306010082.0Jan 2020Jan 2024Sep 2026
Multimodality score, 2020 to 2026
DateIndex
Jan 202012.0
Jul 202015.0
Jan 202121.0
Jul 202127.0
Jan 202234.0
Jul 202241.0
Jan 202348.0
Jul 202355.0
Jan 202463.0
Jul 202469.0
Jan 202574.0
Jul 202578.0
Jan 202680.0
Apr 202680.0
Jul 202681.0
Sep 202682.0