Claim record · edition 1.0.0

si-011Established result

SWE-bench measures whether a system produces a patch that passes a repository’s existing test suite for a real issue.

That is the benchmark definition given by its authors.

Limits of this claim

Passing tests is not correct engineering. The benchmark does not measure design quality, maintainability, review, or deployment safety, and its authors do not claim it does.

Supporting source records · 1

benchmark documentationContent verified

SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Jimenez et al. · 2023

The benchmark definition for test-verified resolution of real repository issues.

Source record →

Related concepts