Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology

What was discovered

AGI Research

Every entry carries its limitations as prominently as its findings, and is mapped to the Index dimensions it bears on. Negative results are indexed with the same weight as positive ones — they are usually more informative.

The archive

All research

Showing 14 of 14 items.

ACLSample

Silent Failure Modes in Long-Context Retrieval

Mid-context retrieval failures present as confident answers rather than detectable errors.

P. Nakamura, A. Lindqvist · Independent

highAGI relevance: high52 citations
CogSciSample

Contamination in Standard Theory-of-Mind Batteries

Canonical social-reasoning tasks appear in training corpora at rates that undermine reported pass rates.

C. Duarte, R. Okonkwo · Independent

mediumAGI relevance: medium40 citations
PreprintSample

Feature-Level Interpretability at Frontier Scale

Analysis methods previously limited to small models operate at frontier scale with partial coverage.

A. Lindqvist, K. Sørensen · Anthropic

highAGI relevance: high103 citations
PreprintSample

Scoring Public AGI Timeline Forecasts, 2015–2022

Capability predictions ran early, deployment predictions ran late, and intervals were universally too narrow.

C. Duarte, D. Almeida · Independent

mediumAGI relevance: medium45 citations