CodeEst. 2023

SWE-bench

SWE-bench evaluates AI models on their ability to resolve real-world GitHub issues from popular open-source Python repositories. Each task requires the model to understand a bug report or feature request, navigate the codebase, and produce a working patch. It tests practical software engineering capabilities far beyond simple code generation, including debugging, testing, and code comprehension at scale.

Metrics

Resolve rate (%) on real GitHub issues

Created By

Carlos E. Jimenez et al. (Princeton)

Top Model Scores

RankModelScoreDate
1Claude Opus 4.662.8%2026-02
2GPT-5.259.3%2026-03
3Gemini 3 Ultra55.7%2026-01
4DeepSeek-V453.2%2026-01
5Grok 449.8%2026-02