Feature-Level Interpretability at Frontier Scale
A. Lindqvist, K. Sørensen — Anthropic
Abstract
We scale feature-level interpretability methods to frontier-size models and report what fraction of computation the resulting features account for. Coverage is partial. We discuss what level of coverage would be required before interpretability could support a decision to remove human review from a deployment, and note that current coverage is far from it.
Key findings
- Methods operate at frontier scale without prohibitive cost
- Coverage is partial; most computation remains uninspected
- Features recovered are stable across training runs
- Coverage sufficient for oversight decisions is not close
Limitations
Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.
- Single model family
- Coverage measurement depends on contested definitions
- No causal validation of the recovered features