Silent Failure Modes in Long-Context Retrieval
P. Nakamura, A. Lindqvist — Independent
Abstract
Long-context evaluation typically measures retrieval accuracy. We instead measure detectability: whether a user without ground truth could identify that retrieval failed. Mid-context failures are found to produce fluent, confident and incorrect output at high rates, making them substantially harder to detect than edge-of-window failures of equivalent frequency.
Key findings
- Mid-context failures produce confident incorrect output at high rates
- Detectability without ground truth is close to chance
- Accuracy metrics systematically understate practical risk
- Effect is stable across model families
Limitations
Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.
- Detectability measured with a limited annotator pool
- Synthetic documents may not reflect real-document structure
- No mitigation is proposed