Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
PreprintSample

Feature-Level Interpretability at Frontier Scale

A. Lindqvist, K. SørensenAnthropic

Abstract

We scale feature-level interpretability methods to frontier-size models and report what fraction of computation the resulting features account for. Coverage is partial. We discuss what level of coverage would be required before interpretability could support a decision to remove human review from a deployment, and note that current coverage is far from it.

Key findings

  • Methods operate at frontier scale without prohibitive cost
  • Coverage is partial; most computation remains uninspected
  • Features recovered are stable across training runs
  • Coverage sufficient for oversight decisions is not close

Limitations

Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.

  • Single model family
  • Coverage measurement depends on contested definitions
  • No causal validation of the recovered features

Continue

Related research

ACLSample

Silent Failure Modes in Long-Context Retrieval

Mid-context retrieval failures present as confident answers rather than detectable errors.

P. Nakamura, A. Lindqvist · Independent

highAGI relevance: high52 citations
CogSciSample

Contamination in Standard Theory-of-Mind Batteries

Canonical social-reasoning tasks appear in training corpora at rates that undermine reported pass rates.

C. Duarte, R. Okonkwo · Independent

mediumAGI relevance: medium40 citations