Concept record · edition 1.0.0
Reliability and compounding error
What the term means here
The tendency of per-step failure probabilities to accumulate across a multi-step trajectory.
Why it matters
It explains why systems that look strong on single-step benchmarks can fail on long tasks, and why error recovery matters more than peak accuracy.
Related claim records · 2
- si-018
METR defines a 50%-task-completion time horizon: the time humans typically take on tasks a model completes with 50% success. Its published measurement reports a horizon of around 50 minutes for one frontier model of that period, and a doubling of the horizon roughly every seven months since 2019, with a possible acceleration in 2024.
- si-019
Increases in measured long-horizon performance are attributed in part to improved reliability and error recovery rather than to reasoning improvements alone.
Sources used here · 2
Measuring AI Ability to Complete Long Software Tasks
Kwa et al. (METR) · 2025
The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.
Source record →AgentBench: Evaluating LLMs as Agents
Liu et al. · 2023
Multi-environment agent evaluation, cited for evaluation design rather than rankings.
Source record →