DeepSeek logo
A

DeepSeek

A Tier · 8.0/10

DeepSeek V4 shipped 2026-04-24: V4-Pro (1.6T/49B active MoE) + V4-Flash (284B/13B active), 1M native context, Hybrid Attention Architecture, open-source on HF. **V4-Pro reached GA 2026-08-13** (Terminal Bench 2.1 87.9, three thinking-effort levels, native Responses API). **Pricing changes 2026-08-16:** peak/off-peak billing raises rates at every hour of the day -- V4-Pro goes from $0.435/$0.87 to $0.66/$1.98 off-peak and $1.32/$3.96 peak

Last updated: 2026-08-13Free tier available

Score Breakdown

7.5
Ease of Use
8.0
Output Quality
9.5
Value
7.0
Features

Benchmark Scores

Benchmarks for DeepSeek V4-Pro (SWE-bench + Arena Elo third-party verified post-launch; knowledge rows are V3.x baseline pending V4 figures)

Chatbot Arena ELOHuman preference rating1220
BenchmarkScore
MMLU90.8%
MMLU-Pro85%
GPQA Diamond79.9%
HumanEval91.5%
SWE-bench80.6%

Last updated: 2026-05-26

Personality & Tone

The open-source reasoning specialist

Tone: Direct and technical. DeepSeek's chat models give compact, math- and code-first answers and are noticeably less chatty than Claude or ChatGPT. When asked to reason, they expose a lot of visible thinking.

Quirks: Refusal patterns differ from Western models -- more permissive on many technical and gray-area prompts, more cautious on China-specific political questions. Community-tuned variants exist with different system prompts and guardrails.

The Good and the Bad

What we like

  • +Pricing is absurdly cheap compared to GPT-4 or Claude -- we're talking 90%+ savings on API calls
  • +DeepSeek-R1 reasoning model genuinely competes with o1 and o3 on math and coding benchmarks
  • +Fully open-source weights mean you can run it locally or fine-tune for your own use case
  • +130M+ users and growing fast, so the ecosystem and community support are solid

What could be better

  • Censorship on politically sensitive topics is real and unavoidable -- it's a Chinese company subject to PRC regulations
  • English output quality is good but noticeably behind Claude or GPT-4 for nuanced writing tasks
  • Hallucinations on niche or domain-specific topics happen more often than with top-tier Western models
  • Service reliability has been spotty during high-demand periods -- the free tier especially suffers from rate limiting

Pricing

Free

$0
  • Web chat access at chat.deepseek.com
  • V4-Flash by default (as of 2026-04-24 launch)
  • Basic usage limits

API -- V4-Flash

$0.14/$0.28/per 1M tokens input/output
  • 284B total / 13B active MoE
  • 1M native context
  • Cheapest frontier-class API on market
  • Pay-as-you-go, no minimum

API -- V4-Pro

$0.435/$0.87/per 1M tokens input/output
  • 1.6T total / 49B active MoE
  • 1M native context
  • GA as DeepSeek-V4-Pro-0813 on 2026-08-13
  • Trails only Gemini 3.1 Pro on world knowledge benchmarks
  • PRICING CHANGES 2026-08-16 16:00 UTC: the $0.435/$0.87 standing rate is replaced by peak/off-peak billing at $0.66/$1.98 off-peak and $1.32/$3.96 peak -- a rise at every hour, not a discount. See knownIssues
  • Still cheap relative to Western frontier models, but the gap narrows sharply after 8/16

Self-hosted (open-source)

$0 + GPU costs
  • MIT license, open weights on HuggingFace
  • V4-Flash is feasible on consumer hardware with quantization
  • V4-Pro needs multi-GPU production infrastructure

System Requirements

Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.

Model variantMinMax
DeepSeek V4-Flash (284B total, 13B active MoE)MIT license, open weights on HuggingFace. Flash is the accessible entry point -- feasible on enthusiast / workstation hardware96 GB RAM + 1× RTX 3090/4090 (Q4 quantization, ~3-5 tok/s)2× H100 FP8 or 1× H200 (FP8 production, fast)
DeepSeek V4-Pro (1.6T total, 49B active MoE)MIT license, open weights. Pro is production multi-GPU territory -- not feasible for individuals512 GB RAM + 4× RTX 4090 (severe quantization, experimental)16× H100 FP8 or 8× H200 (full 1.6T production)
DeepSeek V3.2 (671B total, 37B active MoE) -- prior version, still availableMIT license -- commercial use OK192 GB RAM + 1× RTX 3090/4090 (IQ2_XXS offload, ~2 tok/s)8× H100 FP8 or 4× H200 (full 671B, production)

Known Issues

  • V4-PRO IS FINALLY GA (2026-08-13, vendor change log) -- AND THE ACCOMPANYING 'PEAK/OFF-PEAK' PRICING IS A PRICE **RISE**, NOT A DISCOUNT. THIS IS THE MOST MISREPORTABLE ITEM ON THIS PAGE, SO READ THE NUMBERS. **(1) The GA itself.** DeepSeek's change log states 'The GA release of DeepSeek-V4-Pro has been rolled out on the APP, Web, and API,' model version **DeepSeek-V4-Pro-0813**, calling convention unchanged (`deepseek-v4-pro`). This closes the watch item open since 7/31, when only V4-Flash graduated and DeepSeek said Pro would 'follow soon.' **Vendor-published GA benchmarks:** HLE (without/with tools) **42.7/60.0**, Terminal Bench 2.1 **87.9**, NL2Repo **61.5**, Cybergym **83.3**, DeepSWE **62.7**, Toolathlon-Verified **74.1**, Agents' Last Exam **25.7**, AutomationBench public **31.8**, plus internal sets DSBench-FullStack **71.1** and DSBench-Hard **67.2**. All first-party. **(2) The pricing change -- effective 16:00 UTC on 2026-08-16.** DeepSeek is moving to peak/off-peak billing with 'off-peak prices set at half the peak-hour prices.' **Peak hours are 01:00-04:00 and 06:00-10:00 UTC**; every other hour is off-peak (so ~7 of 24 hours are peak). New V4-Pro rates per 1M: **off-peak $0.022 cache-hit / $0.66 cache-miss input / $1.98 output**; **peak $0.044 / $1.32 / $3.96**. New V4-Flash rates: **off-peak $0.007 / $0.22 / $0.66**; **peak $0.014 / $0.44 / $1.32**. **NOW COMPARE TO WHAT YOU PAY TODAY:** V4-Pro is currently **$0.435 input / $0.87 output**, and V4-Flash **$0.14 / $0.28**. So **even the cheapest new off-peak rate is above the current standing price** -- V4-Pro off-peak input rises ~52% and output ~128%; at peak, input roughly triples and output rises ~355%. V4-Flash off-peak output rises ~136%. **There is no time of day at which the new pricing is cheaper than the old.** Framing it as a discount scheme (which the 'off-peak is half of peak' wording invites) is wrong; the correct read is a substantial across-the-board increase with a time-of-day surcharge layered on top. **This also supersedes our own long-standing note that the 75%-off V4-Pro rate had become the permanent standing price** -- that was true from 2026-05-26 until this change, and it ends on 8/16. **(3) Other GA changes:** thinking effort is now three levels (**low / high / max**) on both Pro and Flash, and the API 'natively supports the OpenAI Responses API format and is specifically adapted for Codex' with a one-click config scriptSource: DeepSeek API change log (api-docs.deepseek.com/updates, entry dated 2026-08-13) and Models & Pricing (api-docs.deepseek.com/quick_start/pricing, carrying both current and 2026-08-16 rate tables) -- both fetched 2026-08-13 · 2026-08-13
  • **[SUPERSEDED 2026-08-13 -- V4-Pro reached GA on that date; see the 2026-08-13 entry above. Kept because it documents that the aggregators calling V4 'GA' in July were wrong at the time.]** V4-FLASH OFFICIAL RELEASE / PUBLIC BETA -- AND V4-PRO IS STILL NOT GA (2026-07-31, vendor changelog; CORRECTS WIDESPREAD AGGREGATOR REPORTING): DeepSeek's own change log states, verbatim: 'The official release of the DeepSeek-V4-Flash API is now in public beta... The DeepSeek-V4-Pro API and the APP/WEB models are unchanged. **The official release of DeepSeek-V4-Pro will follow soon.**' So only FLASH graduated -- aggregators claiming 'DeepSeek V4 went GA in mid/late July' are wrong, and V4-Pro remains Preview. No API change needed: set the model name to `deepseek-v4-flash`. **Vendor-published V4-Flash-0731 benchmarks** (agent-focused, stated as 'far exceeding V4-Pro-Preview'): Terminal Bench 2.1 **82.7**, NL2Repo **54.2**, Cybergym **76.7**, DeepSWE **54.4**, Toolathlon verified **70.3**, Agent Last Exam **25.2**, Automation Bench (public) **25.1**, plus internal sets DSBench-FullStack 68.7 and DSBench-Hard 59.6. Caveats DeepSeek states itself: code-agent numbers were run with the **DeepSeek Harness minimal mode (which it says is still 'to be released soon')** at max effort, topp=0.95, temperature=1.0 -- so they are not straightforwardly reproducible yet, and all figures are first-party. Architecture note: **V4-Flash-0731 keeps the same architecture and size as V4-Flash-Preview and was only re-post-trained.** It also now **natively supports the Responses API format and is specifically adapted for Codex**Source: DeepSeek API change log (api-docs.deepseek.com/updates, fetched 2026-08-03) · 2026-07-31
  • CORRECTION -- THE 2x PEAK-HOUR PRICING HAS NOT ACTUALLY STARTED (re-verified 2026-08-03): our earlier entry described time-of-day pricing as shipping alongside the mid-July V4 release. It has not. DeepSeek's pricing page still frames it in the future tense -- it 'will soon adopt a peak/off-peak pricing policy', with peak hours 09:00-12:00 and 14:00-18:00 Beijing time billed at 2x, and explicitly: '**The effective date will be subject to the official announcement.**' No such announcement has been published as of 2026-08-03. Treat 2x peak pricing as ANNOUNCED-BUT-NOT-IN-EFFECT; current rates are the standing ones shown in the pricing table aboveSource: DeepSeek pricing docs (api-docs.deepseek.com/quick_start/pricing, re-checked 2026-08-03) · 2026-08-03
  • LEGACY API ALIASES RETIRE 2026-07-24 (vendor-primary, HARD deadline): **`deepseek-chat` and `deepseek-reasoner` will be fully retired and inaccessible after July 24, 2026, 15:59 UTC.** The aliases currently route to deepseek-v4-flash (non-thinking/thinking respectively). Migration is a one-line change: keep base_url, update `model` to `deepseek-v4-pro` or `deepseek-v4-flash` -- but note the gotcha that `deepseek-reasoner` maps to FLASH-tier thinking, not V4-Pro, so 'upgrading' to Pro changes both cost and behavior. Any production code still pinned to the legacy aliases breaks on the 24thSource: DeepSeek API docs (api-docs.deepseek.com/news/news260424) · 2026-07-24
  • V4 OFFICIAL RELEASE MID-JULY + FIRST PEAK/OFF-PEAK PRICING (announced 2026-06-30): DeepSeek scheduled the **official (non-preview) V4 release for mid-July 2026**, with 1M context across the lineup -- and will introduce **time-of-day API pricing for the first time: peak hours (9:00-12:00 and 14:00-18:00 daily) billed at 2x the off-peak rate**, effective alongside the release. STATUS as of 2026-07-22: the vendor news page (api-docs.deepseek.com/news/news260424) STILL labels V4 as 'Preview' and the peak/off-peak rate card has not been published as text (only a pricing image), so treat 'GA' as imminent-but-not-confirmed; community reporting points to a WAIC-timed reveal (~7/20-26, Shanghai). What IS locked is the **7/24 15:59 UTC legacy-alias retirement** (see entry above) -- that is the hard, vendor-confirmed date. If you batch heavy workloads, shifting them off-peak will halve token costs once the rate card landsSource: TechNode (technode.com/2026/06/30/deepseek-to-launch-v4-in-mid-july-with-new-peak-time-api-pricing/), DeepSeek API docs (api-docs.deepseek.com/news/news260424, re-checked 2026-07-22 -- still 'Preview') · 2026-06-30
  • Regional availability restrictions: EU, Canada, South Korea, Australia, and India issued formal restrictions or bans on deployment of DeepSeek-V3 and the enterprise API in Q1 2026 over data-residency concerns (traffic routing through mainland China). Germany's BSI confirmed classified metadata leak from a parliamentary pilot. If you're deploying DeepSeek in any of these jurisdictions, check local compliance guidance before shipping; self-hosted open-weights deployment is often the workaround but changes the operational pictureSource: National CSIRT/BSI statements (aggregated), Alibaba policy analysis · 2026-Q1
  • DeepSeek V4 SHIPPED 2026-04-24. Two-model family released simultaneously: V4-Pro (1.6T total / 49B active MoE) and V4-Flash (284B / 13B active MoE). Both default to 1M context natively, use DeepSeek's new Hybrid Attention Architecture, and are open-sourced on HuggingFace under MIT license. V4-Pro trails only Gemini 3.1 Pro on world-knowledge benchmarks per early third-party runs. API pricing: Flash $0.14/$0.28, Pro $1.74/$3.48 per 1M tokens -- still 3-10x cheaper than Western frontier models. Tier-1 coverage: Bloomberg, CNBC, TechCrunch, Simon Willison blog. This closes out the 'V4 imminent' watchlist item that was open since 2026-04-03 Reuters pre-reportSource: DeepSeek API docs, Bloomberg, CNBC, TechCrunch, Simon Willison · 2026-04-24
  • **[SUPERSEDED 2026-08-13 -- this permanent rate ends at 16:00 UTC on 2026-08-16, when peak/off-peak billing raises V4-Pro to $0.66/$1.98 off-peak and $1.32/$3.96 peak. See the 2026-08-13 entry above. Kept because 'DeepSeek made the price war permanent' was accurate for nearly three months and is still widely cited.]** PRICE CUT NOW PERMANENT (confirmed 2026-05-26 via the official pricing page): the 75%-off V4-Pro promo does NOT revert on 2026-05-31. DeepSeek's pricing docs state V4-Pro pricing 'will be officially adjusted to 1/4 of the original price after the 75% discount promotion ends 2026/05/31 15:59 UTC' -- i.e. the discounted rate ($0.435 input / $0.87 output per 1M; cache-hit input $0.003625/M) becomes the new standing list price, down from $1.74 / $3.48. Tech press (The Next Web, Engadget) framed it as DeepSeek making the price war permanent. V4-Flash is unchanged at $0.14 / $0.28. The 'lock in now before the promo ends' framing no longer applies -- this is simply the price nowSource: DeepSeek pricing docs (api-docs.deepseek.com/quick_start/pricing), The Next Web, Engadget · 2026-05-26
  • Third-party verification (T+3 days post-launch): Artificial Analysis Intelligence Index pegs V4-Pro at 52 (#2 open-weight, behind Kimi K2.6) and V4-Flash at 47. Vals AI: V4 is #1 open-weight on Vibe Code Bench 'and it's not close', plus #1 open-weight on SWE-bench. SWE-bench Verified 80.6% (effectively tied with Claude Opus 4.6's 80.8%). Codeforces 3206 surpasses GPT-5.4 (3168) -- highest competitive-programming score at release. GDPval-AA agentic 1554 leads all open-weight models. BUT LMSYS Chatbot Arena Elo around 1220 places V4-Pro alongside GPT-4o and Claude 4 Sonnet, not at the Opus-class frontier (1280+). Simon Willison's pelican-SVG community test produced visibly weak output from V4-Pro (one wing, oversized body) and concluded V4-Pro is 3-6 months behind US frontier labs at a fraction of the cost. Practical verdict: best-in-class open-weight for code/agents/math, mid-pack for general chat quality, weakest for creative/visual generation. Hallucination rate 94%/96% (Pro/Flash) per AA-Omniscience -- caveat for fact-sensitive workloadsSource: Artificial Analysis, Vals AI, Simon Willison, LMSYS Chatbot Arena, Codeforces · 2026-04-27
  • Refuses to engage with questions about Tiananmen Square, Taiwan sovereignty, and other politically sensitive topics per Chinese regulationsSource: Reddit r/LocalLLaMA · 2026-01
  • API latency spikes during peak hours, sometimes timing out entirely on longer reasoning chainsSource: GitHub Issues · 2026-02

Best for

Developers and teams who need strong reasoning and coding capabilities on a budget. If you're building AI features and can't justify GPT-4 API costs, DeepSeek is the obvious first stop.

Not for

Anyone working on content that touches geopolitical topics, or teams that need guaranteed uptime and enterprise SLAs. Also not ideal if your primary use case is creative English writing.

Our Verdict

DeepSeek is the real deal when it comes to bang-for-your-buck AI. The reasoning capabilities are legitimately impressive, and the open-source angle gives it a flexibility that closed models can't match. The censorship limitations are a dealbreaker for some use cases, and the writing quality trails behind Claude and GPT-4. But for coding, math, and analytical tasks? It's hard to argue with near-frontier performance at a fraction of the cost.

Sources

  • DeepSeek API Change Log: V4-Pro GA (DeepSeek-V4-Pro-0813), benchmarks, thinking-effort levels, 2026-08-16 pricing change (2026-08-13) (accessed 2026-08-13)
  • DeepSeek Models & Pricing: current rates plus the peak/off-peak table effective 16:00 UTC 2026-08-16 (accessed 2026-08-13)
  • DeepSeek V4 API launch announcement + 7/24 alias retirement (re-checked 2026-07-22) (accessed 2026-07-22)
  • Bloomberg: DeepSeek unveils newest flagship (2026-04-24) (accessed 2026-04-24)
  • CNBC: DeepSeek V4 LLM preview (2026-04-24) (accessed 2026-04-24)
  • TechCrunch: DeepSeek V4 closes gap with frontier (2026-04-24) (accessed 2026-04-24)
  • Simon Willison: DeepSeek V4 (accessed 2026-04-24)
  • Artificial Analysis: DeepSeek V4 Pro + Flash leading open weights (accessed 2026-04-27)
  • Vals AI: DeepSeek V4-Pro model card (accessed 2026-04-27)
  • DeepSeek pricing docs (75% V4-Pro promo through 2026-05-31) (accessed 2026-04-28)
  • DeepSeek official site (accessed 2026-04-24)
  • Artificial Analysis benchmarks (accessed 2026-04-24)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to DeepSeek

Llama 4 (Meta) logo

Llama 4 (Meta)

Meta's open-weights family -- Scout (10M context), Maverick (multimodal 400B MoE). NOTE: Meta's frontier work moved to the proprietary Muse Spark line in April 2026; Llama remains downloadable and supported but is effectively in maintenance mode

B
7.9/10
Free tierFrom $0
Llama 4 Scout has a 10M token context wi...Llama 4 Maverick is natively multimodal ...
Updated 2026-06-09
Mistral AI logo

Mistral AI

European AI lab with open and commercial models -- Le Chat is now **Vibe** (May 28 2026): one agent across Work Mode + Code Mode with a VS Code extension and CLI, powered by Mistral Medium 3.5 (128B dense, 256k context, 77.6% SWE-Bench Verified). Newest release: **Shieldstral 1.0** (Aug 4 2026), a 3B Apache 2.0 multimodal safety classifier that runs on one 16GB GPU. Earlier 2026 line: Small 4 (119B MoE Apache 2.0), Medium 3, Voxtral TTS

B
7.5/10
Free tierFrom $0
Mistral Medium 3.5 (April 29 2026) is Mi...Vibe Remote Agents (also 4/29) lets you ...
Updated 2026-08-13
Gemma 4 (Google) logo

Gemma 4 (Google)

Google DeepMind's open-weights model family -- multimodal, 256K context, runs on edge devices

A
8.3/10
Free tierFrom $0
Apache 2.0 license -- truly permissive, ...Multimodal: handles text + image input (...
Updated 2026-04-19
DiffusionGemma (Google) logo

DiffusionGemma (Google)

Google DeepMind's experimental open-weights TEXT-DIFFUSION model (June 10, 2026) -- 26B MoE (3.8B active), Apache 2.0, generates 256-token blocks in parallel with bidirectional attention for up to 4x faster output (1,000+ tok/s on H100). Trades some quality vs Gemma 4 for raw speed

C
6.8/10
Free tierFrom $0
Text diffusion instead of autoregression...First open-weights text-diffusion model ...
Updated 2026-06-10
Qwen (Alibaba) logo

Qwen (Alibaba)

Alibaba's open-weights + API family -- Qwen3.8-Max flagship previewed at WAIC (Jul 19 2026: 2.4T sparse-MoE multimodal, closed preview, 'second only to Fable 5'), Qwen 3.7 Max GA (SWE-Bench Pro 60.6%, Terminal-Bench 69.7%, $2.50/$7.50 per 1M), Qwen3.6-27B dense Apache 2.0 (beats the 397B MoE on coding from one consumer GPU)

A
8.8/10
Free tierFrom $0
Qwen 3.6-Plus (launched Mar 30 2026) is ...Qwen3.5 Small (0.8B / 2B / 4B / 9B) is t...
Updated 2026-07-22
GLM / Z.ai (Zhipu AI) logo

GLM / Z.ai (Zhipu AI)

Zhipu AI's open-weights flagship -- GLM-5.2 (launched 2026-06-13) is a ~753B-parameter MoE with a 1M-token context and the new IndexShare sparse-attention architecture (~2.9x lower per-token FLOPs at 1M context), MIT licensed. Vendor benchmarks put SWE-Bench Pro at 62.1 (up from GLM-5.1's 58.4) and it tops the Artificial Analysis open-weights Intelligence Index; VentureBeat reports it beats GPT-5.5 on several long-horizon coding benchmarks at roughly 1/6 the cost. Drop-in for Claude Code / Cline / OpenCode. Still trained outside the Nvidia stack on Huawei Ascend silicon

A
8.0/10
Free tierFrom $0
GLM-5.1 (2026-04-07) topped SWE-Bench Pr...First frontier model trained entirely on...
Updated 2026-07-09
Kimi K3 (Moonshot) logo

Kimi K3 (Moonshot)

Moonshot's 2.8T-parameter Kimi K3 (launched 2026-07-16/17) is the largest open-weight model ever released -- 1M context, multimodal, $3/$15 per 1M via API, ranked best-available on Arena.AI at launch. WEIGHTS SHIPPED ~2026-07-26/27 on Hugging Face (2.8T total / 104B activated, safetensors) under a custom Kimi K3 License, not the Modified MIT of the K2 line

A
8.1/10
Free tierFrom $3 / $15
Frontier-tier performance -- Elo 1309 on...Beats Claude Opus 4.5 on several coding ...
Updated 2026-08-06
Nemotron (Nvidia) logo

Nemotron (Nvidia)

Nvidia's open-weights family -- hybrid Mamba-Transformer MoE architecture, optimized for efficient reasoning on Nvidia hardware. Nemotron 3 Ultra (550B total / 55B active) shipped 2026-06-04 as the family flagship, joining Super (120B/12B, March) and Nano

B
7.8/10
Free tierFrom $0
Hybrid Mamba-Transformer architecture dr...Nemotron 3 Super activates only 12B of 1...
Updated 2026-07-05
MiniMax M3 logo

MiniMax M3

MiniMax's coding/agent flagship -- M3 (June 1 2026): 1M-token context, MSA sparse attention (>15x decoding speedup at long context), SWE-Bench Pro 59.0%, Terminal-Bench 66.0%. OPEN WEIGHTS LIVE on HuggingFace since June 12 (~428B total / ~23B active, native multimodal, minimax-community license)

A
8.4/10
Free tierFrom $0
229B/10B-active MoE delivers Tier-1 agen...Sparse MoE design: ~10B active params du...
Updated 2026-08-03
Falcon (TII) logo

Falcon (TII)

UAE's Technology Innovation Institute open-weights family -- Falcon 3 optimized for efficient sub-10B deployment on consumer hardware

B
7.1/10
Free tierFrom $0
Apache 2.0 license -- fully permissive f...Sub-10B sizes run on any consumer GPU or...
Updated 2026-04-13
gpt-oss (OpenAI) logo

gpt-oss (OpenAI)

OpenAI's FIRST open-weight models -- gpt-oss-120b (single 80GB GPU, near parity with o4-mini on reasoning) and gpt-oss-20b (runs on 16GB edge devices). Apache 2.0. Launched 2025-08-05. gpt-oss-safeguard ships in 2026 as the safety-tuned variant

A
8.1/10
Free tierFrom $0
First-ever OpenAI open-weight release --...gpt-oss-120b approaches o4-mini on reaso...
Updated 2026-04-17
IBM Granite 4.0 logo

IBM Granite 4.0

IBM's enterprise-focused open-weight family -- Granite 4.0 hybrid Mamba-2 + transformer architecture (70-80% memory reduction vs pure transformer), 3B to 32B sizes, Apache 2.0. First open model family to secure ISO 42001 certification. Nano 350M runs on CPU with 8-16GB RAM. 3B Vision variant landed 2026-04-01

A
8.2/10
Free tierFrom $0
Hybrid Mamba-2 + transformer architectur...Granite 4.0 Nano (350M and 1.5B) is genu...
Updated 2026-04-17
Arcee Trinity-Large-Thinking logo

Arcee Trinity-Large-Thinking

Arcee AI's US-made open-weight frontier reasoning model -- launched 2026-04-01. 398B total params, ~13B active. Sparse MoE (256 experts, 4 active = 1.56% routing). Apache 2.0, trained from scratch. #2 on PinchBench trailing only Claude 3.5 Opus. ~96% cheaper than Opus-4.6 on agentic tasks

A
8.1/10
Free tierFrom $0
Rare US-made frontier-tier open-weight r...Trained from scratch (not a fine-tune) a...
Updated 2026-04-17
Olmo 3 (AI2) logo

Olmo 3 (AI2)

Allen Institute for AI's fully-open frontier reasoning models -- Olmo 3 family (2025-11-20) includes 7B and 32B sizes, four variants (Base, Think, Instruct, RLZero). Apache 2.0 with fully open data + checkpoints + training logs. Olmo 3-Think 32B matches Qwen3-32B-Thinking at 6x fewer training tokens

B
7.9/10
Free tierFrom $0
FULLY OPEN is a different category than ...Olmo 3-Think 32B matches Qwen3-32B-Think...
Updated 2026-04-17
AI21 Jamba2 logo

AI21 Jamba2

AI21 Labs' hybrid SSM-Transformer (Mamba-style) open-weight family -- Jamba2 launched 2026-01-08. Two sizes: 3B dense (runs on phones / laptops) and Jamba2 Mini MoE (12B active / 52B total). Apache 2.0, 256K context, mid-trained on 500B tokens

A
8.0/10
Free tierFrom $0
Hybrid SSM-Transformer (Mamba-style) arc...Jamba2 3B dense runs realistically on iP...
Updated 2026-04-17
StepFun Step 3.7 Flash logo

StepFun Step 3.7 Flash

StepFun's (China) agent-focused open-weight family -- Step 3.7 Flash (May 28 2026): 198B sparse MoE vision-language model, ~11B active, 256K context, Apache 2.0, ~400 tok/s, SWE-Bench Pro 56.3. Supersedes Step 3.5 Flash (Feb 2026) as the flagship

B
7.8/10
Free tierFrom $0
Step 3.5 Flash at 196B total / 11B activ...Agent-focused tuning explicitly -- tool ...
Updated 2026-06-10
Cohere Command A logo

Cohere Command A

Cohere's enterprise-multilingual flagship -- 111B params, 256K context, runs on 2x H100. 23 languages. CC-BY-NC 4.0 on weights (research / non-commercial), commercial requires Cohere enterprise contract. Follow-ups: Command A Reasoning + Command A Vision

B
7.5/10
Free tierFrom $0
Best-in-class multilingual open-weight m...Runs on just 2x H100 at FP16 for the ful...
Updated 2026-04-17
LongCat-2.0 (Meituan) logo

LongCat-2.0 (Meituan)

Meituan's open-source 1.6T-parameter MoE (~48B active) with native 1M-token context, MIT license -- trained entirely on domestic Chinese AI ASICs and revealed as the stealth 'Owl Alpha' model that had been topping OpenRouter

B
7.9/10
Free tierFrom $0
Genuinely open frontier scale: 1.6T tota...Native long context: LongCat Sparse Atte...
Updated 2026-07-05
Inkling (Thinking Machines Lab) logo

Inkling (Thinking Machines Lab)

Mira Murati's $12B lab ships its first model (2026-07-15): a 975B/41B-active open-weights MoE that reasons natively over text, images, and audio with a 1M-token context -- positioned not as the strongest model, but as the best starting point for fine-tuning via Tinker

A
8.0/10
Free tierFrom $0
First US frontier-adjacent open-weights ...Natively multimodal INPUT across text, i...
Updated 2026-07-18
Bonsai 27B (PrismML) logo

Bonsai 27B (PrismML)

The first 27B-class model that runs on a phone (2026-07-14) -- ternary and 1-bit quantizations of Qwen3.6 27B squeeze a multimodal, tool-calling, 262K-context model into 3.9-5.9GB under Apache 2.0

B
7.9/10
Free tierFrom $0
Genuine first: a 27B-class model running...Keeps the grown-up capabilities: multi-s...
Updated 2026-07-18