AGI Index dimension · weight 6%
Tool use
Operating external systems to extend capability.
Selecting, invoking and composing external tools — search, execution, APIs, other models — and correctly interpreting what comes back, including failures.
The argument
Evidence and counter-evidence
Both sides are published at equal weight. A framework that only records what raises a score is not measuring anything.
Raises the score
- Tool invocation is reliable enough to sit in production paths at scale, which is a stronger signal than any benchmark.
- Composition across many tools within one task now works, including recovery from tool-level errors.
Holds it down
- Systems under-use tools they should reach for and over-use ones they should not, particularly when the task looks answerable from memory.
- Error interpretation is shallow: a failed call is often retried unchanged.
What the score reads from
Tool-call correctness
Multi-tool composition
Related coverage
Everything bearing on tool use
News
Context
Where this sits against the rest
- Multimodality82
- Coding81
- Reasoning74
- Memory67
- Generalization66
- Planning61
- Learning59
- Autonomy53
- Social intelligence52
- Scientific discovery48
- Embodiment41
How these scores are producedMethodology v1.2