Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Anthropic6 min readSample

Interpretability tooling scales to frontier-size models

Feature-level analysis previously limited to small models is reported working at frontier scale, though coverage remains partial.

highAGI relevance: highalignmentsafety

Agitology analysis

Why it matters

Whether autonomy can safely increase depends on whether internal reasoning can be inspected. Interpretability that scales is a precondition for removing human review, not a nice-to-have.

Key developments

  • Feature-level analysis functioning at frontier scale
  • Coverage partial; most computation remains uninspected
  • Not yet integrated into deployment decisions

Index impact

How this development moved — or failed to move — the dimensions it bears on.

  • +0.3AutonomyEnabling condition rather than capability.

Every Tuesday

The AGI Brief

Five developments, three papers, one forecast, and what moved the Index.

Every Tuesday. No tracking, no sponsor content in the body. Unsubscribe in one click.