Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

AGI Index dimension · weight 6%

Tool use

Operating external systems to extend capability.

Selecting, invoking and composing external tools — search, execution, APIs, other models — and correctly interpreting what comes back, including failures.

The argument

Evidence and counter-evidence

Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.

Raises the score

  • Tool invocation is reliable enough to sit in production paths at scale, which is a stronger signal than any benchmark.
  • Composition across many tools within one task now works, including recovery from tool-level errors.

Holds it down

  • Systems under-use tools they should reach for and over-use ones they should not, particularly when the task looks answerable from memory.
  • Error interpretation is shallow: a failed call is often retried unchanged.

What the score reads from

  • Tool-call correctness

    High and still improving.

  • Multi-tool composition

    Solid within a session; brittle across long chains.