Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
ICLRSample

Reconstructed Held-Out Evaluation of Cross-Domain Transfer

R. Okonkwo, P. NakamuraIndependent

Abstract

Claims of cross-domain transfer in frontier systems rest on evaluations whose contamination status is unknown. We construct held-out equivalents for six widely cited transfer benchmarks using items generated after the relevant training cut-offs and matched for difficulty. Measured transfer falls substantially across all six. We argue that public benchmarks are structurally unable to measure generalization and propose construction criteria for evaluations that can.

Key findings

  • Measured transfer falls substantially on all six reconstructed benchmarks
  • The largest drops occur on the benchmarks most often cited in capability claims
  • Difficulty-matching rules out a naive difficulty confound
  • Proposed construction criteria for contamination-resistant evaluation

Limitations

Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.

  • Item generation is expensive and does not scale to continuous evaluation
  • Difficulty matching is approximate
  • Findings apply to the six benchmarks studied and may not generalise

Continue

Related research

ACLSample

Silent Failure Modes in Long-Context Retrieval

Mid-context retrieval failures present as confident answers rather than detectable errors.

P. Nakamura, A. Lindqvist · Independent

highAGI relevance: high52 citations
CogSciSample

Contamination in Standard Theory-of-Mind Batteries

Canonical social-reasoning tasks appear in training corpora at rates that undermine reported pass rates.

C. Duarte, R. Okonkwo · Independent

mediumAGI relevance: medium40 citations