Concept record · edition 1.0.0
Coding agents
What the term means here
Agent systems operating on real codebases through file, shell, and test interfaces.
Why it matters
The most measured agentic domain, and the clearest demonstration that interface design changes outcomes independently of the model.
Related claim records · 3
- si-010
The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed.
- si-011
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.
Sources used here · 2
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Jimenez et al. · 2023
The benchmark definition for test-verified resolution of real repository issues.
Source record →SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Yang et al. · 2024
The load-bearing source for separating scaffolding capability from model capability.
Source record →