Claim record · edition 1.0.0

si-010Established result

The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed.

This is the explicit finding of the SWE-agent work: interface affordances, not only model quality, determine outcomes on repository tasks.

Limits of this claim

It follows that an agentic benchmark score measures a model-plus-scaffold system. Attributing such a score to the model alone is unsupported, and this map never does so.

Supporting source records · 2

primary paperContent verified

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Yang et al. · 2024

The load-bearing source for separating scaffolding capability from model capability.

Source record →
benchmark documentationContent verified

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Jimenez et al. · 2023

The benchmark definition for test-verified resolution of real repository issues.

Source record →

Related concepts