Concept record · edition 1.0.0

Reliability and compounding error

What the term means here

The tendency of per-step failure probabilities to accumulate across a multi-step trajectory.

Why it matters

It explains why systems that look strong on single-step benchmarks can fail on long tasks, and why error recovery matters more than peak accuracy.

Related claim records · 2

Sources used here · 2

research organizationContent verified

Measuring AI Ability to Complete Long Software Tasks

Kwa et al. (METR) · 2025

The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.

Source record →
benchmark documentationContent verified

AgentBench: Evaluating LLMs as Agents

Liu et al. · 2023

Multi-environment agent evaluation, cited for evaluation design rather than rankings.

Source record →

Related concepts