Inference-Time Compute Scaling Beyond Expected Saturation
Gains from additional inference compute persist across two further orders of magnitude on long-chain problems.
PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means
What was discovered
Every entry carries its limitations as prominently as its findings, and is mapped to the Index dimensions it bears on. Negative results are indexed with the same weight as positive ones — they are usually more informative.
Most AGI-relevant
The archive
Showing 14 of 14 items.
Gains from additional inference compute persist across two further orders of magnitude on long-chain problems.
Reported transfer falls materially when evaluation items are rebuilt to exclude training overlap.
A single policy transfers across structurally different robot platforms with limited degradation.
Continual learning methods effective at small scale fail to hold at frontier model sizes.
Production success rates fall materially below benchmark scores, and the gap is widening.
Mid-context retrieval failures present as confident answers rather than detectable errors.
A propose–run–measure–revise cycle sustained for one month without human intervention.
Canonical social-reasoning tasks appear in training corpora at rates that undermine reported pass rates.
Analysis methods previously limited to small models operate at frontier scale with partial coverage.
Junior postings fall faster than senior postings in occupations with high task-automation exposure.
Measured productivity gains shrink toward zero as correctness requirements tighten.
Intervention rates fall below one per eight-hour run on constrained engineering tasks.
Near-human scene understanding coexists with persistent failure at counting, ordering and localisation.
Capability predictions ran early, deployment predictions ran late, and intervals were universally too narrow.