The oversight question is being settled by commercial pressure, not evidence
Whether human review can be removed is an empirical question. It is being answered by cost curves instead.
The position
Interpretability coverage is nowhere near what an oversight decision requires, and deployment is proceeding on the assumption that it is.
There is a question at the centre of autonomous deployment that I do not think is being asked clearly: what evidence would justify removing a human from a loop where being wrong is expensive?
It is answerable in principle. It has an empirical form. And it is currently being answered in practice by whether supervision is affordable, which is a different question with a different answer.
What oversight assumes
Human review is the safeguard nearly every deployment relies on, and it has a property that is rarely stated: it degrades as the system gets more capable. A reviewer can check work slightly beyond their own competence. Well beyond it, they are not reviewing — they are approving.
So the safeguard weakens exactly as the thing it guards against strengthens. Any deployment plan that treats human review as a fixed control is mispricing it.
Where interpretability actually is
Feature-level analysis now runs at frontier scale, which two years ago it did not. That is real progress and I do not want to undersell it.
But the number that matters is coverage: what fraction of a model's computation these methods account for. It is a small fraction. Not a small fraction that is rapidly growing — a small fraction that is growing at roughly the rate model scale is growing, which means the ratio is not obviously improving.
We can now inspect frontier models. We cannot yet inspect enough of one to know whether we have found the part that would hurt us.
Why this is being decided badly
Supervision is the dominant cost of autonomous deployment. As capability rises, the commercial case for removing it strengthens continuously. The safety case does not — it requires a discrete demonstration that has not been made.
So the threshold gets crossed gradually, in individual product decisions, each defensible on its own, none of which is the decision to remove oversight. That is how consequential thresholds are usually crossed: not by a choice, but by an accumulation.
What would change my position
- Interpretability coverage measured, published and improving faster than model scale.
- Calibration that holds under distribution shift — systems that know they are outside their competence, tested where it matters and not only where it is cheap.
- Published intervention rates from consequential deployments, including the ones that went badly.
- A demonstrated case of a system halting on a failure its operators had not anticipated. One would be more informative than a year of benchmarks.
The Index scores interpretability progress as an enabling condition rather than a capability, and I think that framing is right. Autonomy at 53 is a capability statement. Whether we should act on it is a different question, and the evidence that would answer it does not exist yet.