Coding

SWE-bench Verified: 2026 AI Leaderboard

Fix real GitHub issues in 12 open-source Python repos.

What it tests

SWE-bench Verified is a 500-issue subset of SWE-bench that has been human-validated as solvable. Each task is a real Python GitHub issue; the model is given the repo, the issue, and must produce a patch that makes the project's test suite pass.

How it is scored

Percentage of issues where the generated patch passes all hidden tests. This is end-to-end agentic coding, not just code-completion. Scores above 70% are state-of-the-art; a year ago it was 30%.

Why it matters

SWE-bench Verified is the closest industry-standard benchmark to 'can this model actually do my job'. It rewards code-reading, multi-file editing, and test-driven iteration -- not just autocomplete.

Leaderboard (6 models)

Sorted by SWE-bench Verifiedscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.

#ModelTierSWE-bench Verified score
1Claude (Anthropic)
Claude Opus 5 (launched 2026-07-24) is the current flagship and the default on Max -- Anthropic released COMPARATIVE claims only, no absolute scores: more than doubles Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, 3x the next-best model on ARC-AGI 3, ~1.5x on Zapier AutomationBench, and beats Fable 5's best OSWorld 2.0 result at just over a third of the cost. Fable 5 (2026-06-09) holds vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%), #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending third-party suites for Opus 5
A80.8%
2DeepSeek
DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures)
A80.6%
3Mistral AI
Mistral Medium 3.5 (vendor-published; third-party verification pending)
B77.6%
4Codex (OpenAI)
GPT-5.2-Codex (HISTORICAL -- this model was RETIRED 2026-07-23; scores retained for trend context only. Current Codex default is GPT-5.6 Sol, whose first-party Codex benchmarks are pending)
A72%
5ChatGPT
GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification)
A72%
6Qwen (Alibaba)
Qwen3.5-397B MoE
A69.4%

About SWE-bench Verified

Creator
Princeton & OpenAI, 2023 (Verified subset 2024)
Unit
% (max 100)

Other benchmarks