Claim record · edition 1.0.0

si-019Active research

Increases in measured long-horizon performance are attributed in part to improved reliability and error recovery rather than to reasoning improvements alone.

The METR paper attributes the horizon increase primarily to greater reliability and ability to adapt to mistakes, combined with logical reasoning and tool use.

Limits of this claim

This is an attribution offered by one research group over one task suite. The decomposition between reliability, reasoning, and scaffolding is not independently established.

Supporting source records · 1

research organizationContent verified

Measuring AI Ability to Complete Long Software Tasks

Kwa et al. (METR) · 2025

The primary definition of the 50%-task-completion time horizon. Abstract read during this pass; it is the only verified source for any time-horizon figure in this map.

Source record →

Related concepts