Concept record · edition 1.0.0
Long-horizon task completion
What the term means here
The ability to carry a task requiring many dependent steps to completion, measured here by the time a human would take on tasks a system completes at a given success rate.
Why it matters
It is the capability most relevant to autonomy claims and the one where measurement is least mature.
What this does not establish
That a measured horizon transfers to real-world software work. The source paper explicitly frames external validity as a limitation and its forward statement as conditional.
Related claim records · 2
- si-018
METR defines a 50%-task-completion time horizon: the time humans typically take on tasks a model completes with 50% success. Its published measurement reports a horizon of around 50 minutes for one frontier model of that period, and a doubling of the horizon roughly every seven months since 2019, with a possible acceleration in 2024.
- si-019
Increases in measured long-horizon performance are attributed in part to improved reliability and error recovery rather than to reasoning improvements alone.
Sources used here · 1
Measuring AI Ability to Complete Long Software Tasks
Kwa et al. (METR) · 2025
The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.
Source record →