Claim record · edition 1.0.0
si-011Established result
SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.
That is the benchmark definition given by its authors.
Limits of this claim
Passing tests is not correct engineering. The benchmark does not measure design quality, maintainability, review, or deployment safety, and its authors do not claim it does.
Supporting source records · 1
benchmark documentationContent verified
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Jimenez et al. · 2023
The benchmark definition for test-verified resolution of real repository issues.
Source record →