Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
ACLSample

Silent Failure Modes in Long-Context Retrieval

P. Nakamura, A. LindqvistIndependent

Abstract

Long-context evaluation typically measures retrieval accuracy. We instead measure detectability: whether a user without ground truth could identify that retrieval failed. Mid-context failures are found to produce fluent, confident and incorrect output at high rates, making them substantially harder to detect than edge-of-window failures of equivalent frequency.

Key findings

  • Mid-context failures produce confident incorrect output at high rates
  • Detectability without ground truth is close to chance
  • Accuracy metrics systematically understate practical risk
  • Effect is stable across model families

Limitations

Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.

  • Detectability measured with a limited annotator pool
  • Synthetic documents may not reflect real-document structure
  • No mitigation is proposed

Continue

Related research