Claim record · edition 1.0.0
si-010Established result
The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed.
This is the explicit finding of the SWE-agent work: interface affordances, not only model quality, determine outcomes on repository tasks.
Limits of this claim
It follows that an agentic benchmark score measures a model-plus-scaffold system. Attributing such a score to the model alone is unsupported, and this map never does so.
Supporting source records · 2
primary paperContent verified
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Yang et al. · 2024
The load-bearing source for separating scaffolding capability from model capability.
Source record →benchmark documentationContent verified
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Jimenez et al. · 2023
The benchmark definition for test-verified resolution of real repository issues.
Source record →