Deliberation budgets scale past the point of diminishing returns
A frontier laboratory reports that allocating substantially more inference compute to hard problems continues to yield gains well beyond where the curve was expected to flatten.
PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means
What happened
Every item is filtered for one question: does it change what we believe about progress toward general intelligence? Items that do carry an explicit Index impact. Items that do not are still worth reading, and say so.
Showing 18 of 18 items.
A frontier laboratory reports that allocating substantially more inference compute to hard problems continues to yield gains well beyond where the curve was expected to flatten.
An independent audit finds that a meaningful share of reported cross-domain transfer disappears once evaluation items are reconstructed to exclude training overlap.
Agent deployments report sustained unsupervised operation on well-specified engineering work, with intervention rates falling below one per eight-hour run.
A general-purpose policy trained on one robot platform is reported to transfer to structurally different hardware with limited degradation.
An autonomous experimentation platform reports a sustained propose–run–measure–revise cycle in a narrow materials domain.
Methods reported to avoid catastrophic forgetting in smaller models are found not to hold at frontier scale.
Draft guidance would require documented human review of consequential decisions taken by autonomous systems in finance and healthcare.
A survey of engineering organisations reports that reviewing machine-written code now consumes more time than the writing it replaced.
Evaluation work finds mid-context retrieval failures that produce confident, fluent and wrong answers rather than visible errors.
Grid interconnection timelines rather than chip supply are reported as the limiting factor on several planned training clusters.
The canonical task batteries used to assess social reasoning appear in training corpora at high rates, undermining reported pass rates.
Publicly released weights are reported within a narrow band of closed frontier systems on deliberation-heavy evaluations.
Feature-level analysis previously limited to small models is reported working at frontier scale, though coverage remains partial.
Labour-market data shows junior-role postings falling faster than senior ones in occupations with high task-automation exposure.
Systems near human baseline on scene understanding continue to fail at counting, ordering and precise localisation.
A retrospective scores public AGI timeline predictions made between 2015 and 2022 against what actually occurred.
Systems scoring highly on agentic evaluations show substantially lower success rates in production environments.
Public compute allocations let academic groups run evaluations previously feasible only inside frontier laboratories.