Inference-Time Compute Scaling Beyond Expected Saturation
R. Okonkwo, D. Almeida, K. Sørensen — Independent
Abstract
The relationship between inference-time computation and task accuracy has been assumed to saturate within roughly an order of magnitude of current deployment budgets. We evaluate that assumption across problem classes with varying dependent-chain length and find that returns persist substantially further than predicted, with the effect concentrated on problems requiring long sequences of dependent inference. We characterise the regime in which additional deliberation continues to pay and the regime in which it does not.
Key findings
- Returns persist across two further orders of magnitude on long-chain problems
- No corresponding effect on problems solvable in few dependent steps
- Cost per solved problem rises faster than accuracy, bounding practical use
- The effect is architecture-independent across the systems tested
Limitations
Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.
- Evaluation set is heavily weighted toward mathematics and formal reasoning
- Contamination cannot be excluded for the public portion of the set
- Economic viability of the high-compute regime is not assessed