Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
PreprintSample

Unsupervised Operating Windows in Production Agent Deployments

J. Halvorsen, T. IshikawaIndependent

Abstract

We instrument production agent deployments and measure the interval between human interventions. On well-specified engineering tasks the rate falls below one intervention per eight-hour run. On open-ended tasks it remains an order of magnitude higher. We separate interventions that correct errors from those that redirect goals.

Key findings

  • Below one intervention per eight-hour run on constrained tasks
  • Open-ended tasks remain an order of magnitude worse
  • Self-halting on detected failure now exceeds silent continuation
  • Goal-redirection interventions are not falling

Limitations

Published at equal prominence to the findings. A paper’s limitations are usually the part that determines how much its result should move your beliefs.

  • Deployments are self-selected and well-resourced
  • Task specification quality is not controlled
  • Intervention classification involved judgement

Continue

Related research