Reasoning

GPQA Diamond: 2026 AI Leaderboard

Graduate-level physics, biology, and chemistry written to defeat Google-search.

What it tests

GPQA (Graduate-level Google-Proof Q&A) Diamond is the hardest subset of a 448-question multiple-choice set written by PhDs in physics, biology, and chemistry. Questions are deliberately designed so that searching the web does not yield the answer.

How it is scored

Four-choice accuracy. Domain PhDs with unlimited internet access score about 65%; non-expert humans with search score roughly 34%. Frontier models in 2026 are hitting the 80s and 90s -- a major inflection.

Why it matters

GPQA Diamond is the most cited reasoning benchmark for frontier LLMs precisely because it resists memorization. A high score implies the model can synthesize knowledge, not just recite training data.

Leaderboard (11 models)

Sorted by GPQA Diamondscore. Tier column shows the tool's overall AIToolTier rank, which blends this benchmark with pricing, features, and real-world usability.

#ModelTierGPQA Diamond score
1ChatGPT
GPT-5.5 (launched 2026-04-23; scores below are the GPT-5.4 baseline -- GPT-5.5 launch benchmarks per OpenAI are logged in Known Issues, pending third-party verification)
A92.8%
2Claude (Anthropic)
Claude Opus 5 (launched 2026-07-24) is the current flagship and the default on Max -- Anthropic released COMPARATIVE claims only, no absolute scores: more than doubles Opus 4.8 on Frontier-Bench, within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, 3x the next-best model on ARC-AGI 3, ~1.5x on Zapier AutomationBench, and beats Fable 5's best OSWorld 2.0 result at just over a third of the cost. Fable 5 (2026-06-09) holds vendor SWE-Bench Pro 80.3% (vs GPT-5.5 58.6%), #1 LMArena Elo 1510 and #1 Artificial Analysis Index 65 as of 6/11. Sonnet 5 (2026-06-30): OSWorld-Verified 78.5%. Legacy Opus-line reasoning-suite scores shown below as baseline pending third-party suites for Opus 5
A91.3%
3GLM / Z.ai (Zhipu AI)
GLM-5.2 (~753B MoE, launched 2026-06-13) -- vendor-published; third-party verification still settling
A91.2%
4Muse Spark (Meta)
Muse Spark
A86%
5Grok
Grok 4.20 (baseline -- Grok 4.5 launched 2026-07-08; vendor-reported 4.5 scores in Known Issues pending third-party verification)
B85%
6Gemma 4 (Google)
Gemma 4 31B
A84.3%
7DeepSeek
DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures). V4.1-Flash (2026-09-10) vendor-only so far: GPQA Diamond 90.9, DeepSWE 74.2, Terminal-Bench 2.1 90.6 -- third-party verification pending
A79.9%
8Qwen (Alibaba)
Qwen3.5-397B MoE
A78.2%
9Nemotron (Nvidia)
Llama-Nemotron Ultra 253B (prior gen -- Nemotron 3 Ultra 550B third-party scores pending)
B70.5%
10Llama 4 (Meta)
Llama 4 Maverick (17B/400B MoE)
B69.8%
11Falcon (TII)
Falcon 3 10B
B42.5%

About GPQA Diamond

Creator
Rein et al., 2023 (NYU/Cohere/Anthropic)
Unit
% (max 100)

Other benchmarks