Scientific discovery: everything except the part that matters
Autonomous laboratories now close the experimental loop. They have not yet decided what is worth investigating, and that is the whole dimension.
Scientific discovery scores 48, the second-lowest dimension in the Index, on a record that includes genuine and consequential results. Structural biology was changed permanently. Materials screening produces candidates that survive experimental validation at rates that alter research economics. A materials platform this year ran a closed experimental loop for a month without human intervention.
All of that is real, and the score is still 48, because the dimension is not scoring acceleration. It is scoring discovery.
The distinction
Every result above shares a structure: a human identified a question worth answering, framed it in a form a system could search over, and the system searched faster and better than people could. That is acceleration within a human-framed programme, and it is enormously valuable.
What no system has yet done is look at a field and identify that a particular question is worth asking — the step that distinguishes a research programme from a search procedure, and the one scientists spend careers learning.
A system that answers questions faster than any human is not doing science if a human is still choosing the questions. It is doing very good instrumentation.
The physical bound
The second constraint is less philosophical. Experiments take the time they take. A system proposing a thousand candidates an hour into a laboratory that can test forty a week has not accelerated anything; it has moved the queue.
This is why the autonomous laboratory result matters more than the proposal-generation results that preceded it. Closing the loop attacks throughput rather than proposal volume, and throughput is the binding constraint.
What would move this score
- A system that identifies a research question its operators had not considered, and the question turns out to be worth pursuing — the single result that would move this dimension the most.
- Independent replication of AI-originated findings at scale. The literature currently has a strong bias toward successes, which makes the base rate unknown.
- Closed-loop platforms in a second and third domain, demonstrating the approach is not specific to materials.
- Any negative result published with the same prominence as the positive ones.
The dimension carries a low confidence rating, and it should. The evidence is thin, the successes are highly visible, and the failures are not published at all.