Source record · content verified
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
benchmark documentation2023
Why this source is here
The benchmark definition for test-verified resolution of real repository issues.
- Verification
- Content-verified on 2026-07-27: the canonical source and its title were resolved during the Atlas review. This is not an endorsement of the source’s argument.
- Authors
- Jimenez et al.
- Identifier
- arXiv:2310.06770
Claims citing this source · 3
- si-010
The design of the agent-computer interface materially changes measured coding-agent performance holding the underlying model fixed.
- si-011
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
- si-023
A benchmark score is a measurement of a system on a task distribution under a scaffolding configuration at a date. It is not a measurement of general capability.