Reconstructed Held-Out Evaluation of Cross-Domain Transfer
R. Okonkwo, P. Nakamura — Independent
Abstract
Claims of cross-domain transfer in frontier systems rest on evaluations whose contamination status is unknown. We construct held-out equivalents for six widely cited transfer benchmarks using items generated after the relevant training cut-offs and matched for difficulty. Measured transfer falls substantially across all six. We argue that public benchmarks are structurally unable to measure generalization and propose construction criteria for evaluations that can.
Key findings
- Measured transfer falls substantially on all six reconstructed benchmarks
- The largest drops occur on the benchmarks most often cited in capability claims
- Difficulty-matching rules out a naive difficulty confound
- Proposed construction criteria for contamination-resistant evaluation
Limitations
Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.
- Item generation is expensive and does not scale to continuous evaluation
- Difficulty matching is approximate
- Findings apply to the six benchmarks studied and may not generalise