Coding

HumanEval: 2026 AI Leaderboard

164 Python programming problems: does the generated code pass unit tests?

What it tests

HumanEval is 164 handwritten Python programming problems with hidden unit tests. A model sees the function signature plus docstring and must generate a body that passes every test.

How it is scored

pass@1 -- the percentage of problems solved on the first attempt. Frontier models in 2026 sit in the 94-99% range, so this benchmark is effectively saturated for top-tier LLMs.

Why it matters

Still useful as a floor-check: any serious coding model should clear 90% here. For real-world discrimination, SWE-bench Verified and LiveCodeBench are the benchmarks that still separate the field.

Leaderboard (13 models)

Sorted by HumanEvalscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.

#ModelTierHumanEval score
1Codex (OpenAI)
GPT-5.2-Codex (launched 2026-04-23 -- SOTA on SWE-Bench Pro and Terminal-Bench 2.0; first-party scores below pending detailed third-party verification)
A95%
2ChatGPT
GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification)
A95%
3Claude (Anthropic)
Claude Fable 5 (launched 2026-06-09; suspended 2026-06-12 by US gov order; RESTORED globally 2026-07-01 after controls lifted 6/30) -- vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%); #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. New default Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending full third-party suites
A94%
4Qwen (Alibaba)
Qwen3.5-397B MoE
A92.5%
5Mistral AI
Mistral Medium 3.5 (vendor-published; third-party verification pending)
B92%
6DeepSeek
DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures)
A91.5%
7Muse Spark (Meta)
Muse Spark
A91%
8Grok
Grok 4.20 (baseline -- Grok 4.5 launched 2026-07-08; vendor-reported 4.5 scores in Known Issues pending third-party verification)
B90%
9Nemotron (Nvidia)
Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending)
B89.6%
10GLM / Z.ai (Zhipu AI)
GLM-5.2 (~753B MoE, launched 2026-06-13) -- vendor-published; third-party verification still settling
A89.1%
11Llama 4 (Meta)
Llama 4 Maverick (17B/400B MoE)
B88%
12Gemma 4 (Google)
Gemma 4 31B
A85%
13Falcon (TII)
Falcon 3 10B
B73.8%

About HumanEval

Creator
OpenAI, 2021
Unit
% (max 100)

Other benchmarks