Kimi K3 (Moonshot) logo
A

Kimi K3 (Moonshot)

A Tier · 8.1/10

Moonshot's 2.8T-parameter Kimi K3 (launched 2026-07-16/17) is the largest open-weight model ever released -- 1M context, multimodal, $3/$15 per 1M via API, ranked best-available on Arena.AI at launch. WEIGHTS SHIPPED ~2026-07-26/27 on Hugging Face (2.8T total / 104B activated, safetensors) under a custom Kimi K3 License, not the Modified MIT of the K2 line

Last updated: 2026-09-05Free tier available

Score Breakdown

6.0
Ease of Use
9.0
Output Quality
8.5
Value
9.0
Features

Benchmark Scores

Benchmarks for Kimi K3 (2.8T, launched 2026-07-16) -- Arena.AI ranked it best-available at launch; vendor claims parity with Fable 5, third-party suites pending. Scores below are K2.6/K2.5-era baselines retained until K3 third-party runs publish

BenchmarkScore
SWE-Bench Pro58.6%
MMLU-Pro (K2.5 baseline)84.8%
GPQA Diamond (K2.5 baseline)80.5%
AIME 2025 (K2.5 baseline)91.2%
LiveCodeBench (K2.5 baseline)74.1%

Last updated: 2026-04-27

Personality & Tone

The long-context note-taker

Tone: Careful and document-focused. Kimi K2.5 shines when you dump a long document in -- replies read as summary-and-citation rather than open chat, leaning on the source material rather than the model's opinions.

Quirks: Context handling is the whole pitch. Without a document to anchor to, replies feel plainer than Qwen or DeepSeek. Native Chinese quality is very strong; English is decent but not class-leading.

The Good and the Bad

What we like

  • +Frontier-tier performance -- Elo 1309 on GDPval-AA, behind only OpenAI and Anthropic flagships
  • +Beats Claude Opus 4.5 on several coding benchmarks per community testing
  • +Unified thinking + non-thinking modes in one model (no need to swap)
  • +256K context window handles large codebases for agentic coding
  • +Weights are genuinely downloadable -- K2.6/K2.7-Code under Modified MIT, and K3 under its own Kimi K3 License (check its terms before commercial use)
  • +Native tool-use and agentic planning trained in -- not bolted on

What could be better

  • Self-hosting is datacenter-only -- K3 is 2.8T params (104B activated) and even the older K2 line needs 4+ H100-class GPUs
  • Moonshot is a smaller lab than DeepSeek/Alibaba -- less Western infrastructure support
  • API pricing ($0.60 in / $3.00 out) is higher than DeepSeek V3.2 ($0.28 in / $0.42 out)
  • PRC content filters apply (Tiananmen, Taiwan, etc.)
  • Documentation is heavily Chinese-first -- English docs trail releases

Pricing

API (Kimi K3)

$3 / $15/per 1M tokens (input/output)
  • K3 (launched 2026-07-16): 2.8T-param open-weight multimodal reasoning model
  • 1M token context window
  • Reasoning effort currently supports only 'max'
  • Live on kimi.com, Kimi app, Moonshot API, and OpenRouter; capacity-limited at launch (frequent 429s)

Self-hosted (Free -- K2.6/K2.7 line)

$0
  • K3 WEIGHTS PUBLISHED ~2026-07-26/27 at huggingface.co/moonshotai/Kimi-K3 -- safetensors, 2.8T total params, 104B activated, 1,048,576-token context
  • K3 ships under its own **Kimi K3 License**, NOT the Modified MIT used for K2 -- read the license terms before commercial use
  • K2.6 + K2.7-Code weights remain on Hugging Face under Modified MIT
  • Fine-tuning permitted

API (Moonshot direct, K2.6)

$0.60/per 1M input tokens
  • K2.6: $0.60 in / $2.50 out (Moonshot direct)
  • 256K context
  • Native video input (mp4/mov/avi/webm)

System Requirements

Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.

Model variantMinMax
Kimi K2.5 (1T total, 32B active MoE)Practically a hosted-only model for most users -- self-hosting requires enterprise hardware256 GB unified RAM Mac Studio M3 Ultra (Q2, ~3 tok/s)8× H200 141 GB FP8 or 16× H100 (production-grade)

Known Issues

  • KIMI K2.7 CODE IS BEING REMOVED FROM GITHUB COPILOT ON 2026-10-02 (announced 2026-09-03, vendor-primary -- GitHub, not Moonshot): GitHub listed **Kimi K2.7 Code** on a four-model deprecation slate to be removed '**across all GitHub Copilot experiences (including Copilot Chat, inline edits, ask and agent modes, and code completions) on October 2nd, 2026**', with **Kimi K3 named as the suggested alternative**. The others on the slate are Gemini 3.5 Flash, Gemini 3.6 Flash and Claude Opus 4.7. **SCOPE, BECAUSE THIS WILL BE MISREPORTED: this is a GitHub distribution decision, not a Moonshot retirement.** Nothing here says K2.7 Code is being discontinued by Moonshot, and Moonshot has published no corresponding notice; the model simply stops being selectable inside Copilot. Direct API and other surfaces are unaffected by this announcement. **THE PRACTICAL TRAP FOR COPILOT ORGS IS THE MIGRATION TARGET, NOT THE REMOVAL.** GitHub warns that '**Copilot Enterprise and Copilot Business administrators may need to enable access to the alternative models through their model policies**' -- and K3, being **open-weight**, is in the category GitHub excludes from default enablement, so it is one of the models most likely to be sitting off. **An org relying on K2.7 Code inside Copilot should verify the K3 policy is enabled before October 2, or it will lose the Kimi line entirely on that date rather than migrating to it.** This is the same open-weight exclusion this page has tracked since the 8/26 default-enablement rollout.Source: GitHub Changelog (github.blog/changelog/2026-09-03-upcoming-deprecation-of-selected-github-copilot-models/) -- fetched 2026-09-05 · 2026-09-03
  • K3 BREAKS INTO WESTERN PRODUCT DISTRIBUTION IN A SINGLE WEEK -- GITHUB COPILOT AND PERPLEXITY (2026-08-04 and 2026-08-06, both vendor-primary): three weeks after launch, K3 stopped being an API-and-weights story and became a model you can select inside mainstream Western products. **GITHUB COPILOT (2026-08-06):** K3 is **generally available** in the Copilot model picker across **Pro, Pro+, Max, Business and Enterprise** -- including base Pro, a broader entitlement than Claude Opus 5 got on 7/24. GitHub's own assessment: it '**shows frontier-level abilities on agentic coding with highly cost-effective pricing**.' Crucially, **GitHub hosts K3 on Fireworks AI**, so requests do not go to Moonshot's API. Billed at provider list pricing under usage-based billing; available in VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, JetBrains, Xcode, Eclipse and GitHub Mobile. It is **off by default for Business and Enterprise** and needs an explicit admin policy, and GitHub pointedly tells admins to 'review open-weight models against their own security, compliance, and data-governance requirements before enabling them.' **PERPLEXITY (2026-08-04):** K3 is available in Perplexity and Perplexity Computer for **Pro and Max** subscribers, '**hosted exclusively on US-based servers**' -- Perplexity's pitch is explicitly '2.8T-parameter mixture-of-experts model with a 1M-token context window and native vision, **without data leaving US infrastructure**.' **WHY THIS MATTERS MORE THAN A NORMAL INTEGRATION:** the single biggest practical objection to Kimi in Western enterprises was jurisdictional, not technical. Both integrations answer it the same way -- re-host the open weights on US infrastructure and sell the model without the Chinese API. That is only possible *because* the weights are public, which makes this the clearest commercial payoff yet for Moonshot's open-weight strategy. It does not, however, resolve the licence question: the **Kimi K3 License** still governs the weights themselves, so re-hosting by a vendor is not the same as your own right to deploy themSource: GitHub changelog (github.blog/changelog/2026-08-06-kimi-k3-is-now-available-in-github-copilot/, fetched 2026-08-06), Perplexity changelog (perplexity.ai/changelog/shared-workspaces-personal-computer-for-windows-and-model-council, dated 08/04/26, fetched 2026-08-06) · 2026-08-06
  • OPEN WEIGHTS SHIPPED (~2026-07-26/27 -- Moonshot announced 7/27 and reporting puts the actual drop a day earlier, so treat the exact day as approximate): Moonshot published **Kimi K3's weights** to Hugging Face at `moonshotai/Kimi-K3`, resolving the open question this page carried since launch. The model card confirms **2.8T total parameters with 104B activated** (the ~50B active figure that circulated in aggregator coverage was wrong), a **1,048,576-token context window**, and safetensors in F32 / BF16 / U8. **License is the bespoke 'Kimi K3 License', NOT the Modified MIT that covers K2.6 and K2.7-Code** -- the card states 'Both the code repository and the model weights are released under the Kimi K3 License', so anyone planning commercial use or a derivative needs to read those terms rather than assuming K2 permissions carry over. Practical effect: K3 becomes the largest open-weight model actually downloadable, though at 2.8T params it is a datacenter-class deployment, not a local oneSource: Hugging Face model card (huggingface.co/moonshotai/Kimi-K3, fetched 2026-07-29) · 2026-07-27
  • MODEL LAUNCH -- KIMI K3 (2026-07-16/17): Moonshot shipped **Kimi K3**, billed as the largest open-weight model ever announced -- **2.8T total parameters**, multimodal reasoning, **1M token context**, live immediately on kimi.com, the Kimi app, the Moonshot API, and OpenRouter at **$3/M input, $15/M output** (vs Fable 5's $50/M output). Unveiled at the World AI Conference in Shanghai 7/17 after appearing on platforms 7/16; the stealth Arena model 'Kivine' was K3 in testing, and Arena.AI ranked it the best available model at launch. Vendor claims it performs competitively with Claude Fable 5 and 'substantially outperforms' Opus 4.8 and GPT-5.6 Sol -- third-party verification pending; treat vendor benchmark claims accordingly. Market reaction was dubbed a 'second DeepSeek shock': TSMC fell 7%, SoftBank 9%, Nasdaq 100 ~1% on 7/17. CAVEATS (updated 2026-07-29 -- (a) and (d) are now RESOLVED, see the weights-release entry above): (a) weights shipped ~2026-07-26/27; (b) API capacity is limited at launch -- OpenRouter flags frequent 429 errors; (c) reasoning effort currently supports only 'max'; (d) the aggregator-reported ~50B active-parameter figure was WRONG -- the model card publishes **104B activated**Source: OpenRouter (openrouter.ai/moonshotai/kimi-k3), Fortune (2026-07-17), Reuters, CNBC, r/LocalLLaMA · 2026-07-17
  • MODEL LAUNCH (2026-06-12): **Kimi K2.7-Code** -- Moonshot's code-specialized variant dropped on HuggingFace (moonshotai/Kimi-K2.7-Code, HN front page 333 points). Specs: 1T-param MoE with 32B active (384 experts), 256K context, Modified MIT license, MoonViT 400M vision encoder, built on K2.6, ~30% fewer thinking tokens than K2.6, forces thinking mode on. API via platform.moonshot.ai (OpenAI- and Anthropic-compatible endpoints). Notably honest self-published benchmarks show it TRAILING the frontier: Kimi Code Bench v2 62.0 vs GPT-5.5's 69.0 and Opus 4.8's 67.4 -- Moonshot is positioning on open-weights value, not SOTA claims. API pricing circulating in aggregators (~$0.19 cached/$0.95 in/$4.00 out per 1M) -- verify on the vendor pricing page before relying on itSource: HuggingFace (huggingface.co/moonshotai/Kimi-K2.7-Code), Hacker News, platform.moonshot.ai · 2026-06-12
  • API DEPRECATION (2026-05-25, vendor docs verbatim: 'The kimi-k2 series models were officially discontinued on May 25, 2026'): retired model ids -- kimi-k2-0905-preview, kimi-k2-0711-preview, kimi-k2-turbo-preview, kimi-k2-thinking, kimi-k2-thinking-turbo. If your code pins any of these, requests now fail; migrate to kimi-k2.6 (or kimi-k2.5, which remains an active model alongside it). K2.6 detail confirmed on the vendor blog: open weights on HuggingFace (moonshotai/Kimi-K2.6), 256K context (262,144 default), agent swarm scaling to **300 sub-agents / 4,000 coordinated steps** (up from K2.5's 100/1,500)Source: Moonshot platform docs (platform.kimi.ai/docs/models), kimi.com/blog/kimi-k2-6, HuggingFace moonshotai/Kimi-K2.6 · 2026-05-25
  • SUPERSEDED (2026-07-16): the June-era 'K3 never shipped / treat K3 claims as fabrication' guidance no longer holds -- Kimi K3 is real and launched July 16-17, 2026 (see the K3 launch entry above). The May Manifold window did resolve NO (K3 missed May by six weeks), and the skepticism was correct at the time; the launch simply came later than the rumor mill claimed. Retained for history: the K2-series API deprecation (5/25) and K2.6 investment preceded K3 rather than replacing it.Source: kimi.com, OpenRouter (openrouter.ai/moonshotai/kimi-k3), Fortune (2026-07-17) · 2026-07-16
  • Kimi K2.6 (GA 2026-04-20) supersedes K2.5 -- 1T total / 32B active MoE, 256K context, adds native video input (mp4/mov/avi/webm). Scores 54 on Artificial Analysis Intelligence Index v4.0, ranked #1 open-weights and #4 overall (three points behind Claude Opus 4.7 / Gemini 3.1 Pro / OpenAI flagships at 57). SWE-Bench Pro 58.6%. Modified MIT license unchanged. Moonshot direct API: $0.60 in / $2.50 out per 1M tokens. OpenRouter blended: ~$0.95 in / $4.00 out. If you were on K2.5, the upgrade is non-breaking on the API side -- Moonshot routes the K2.6 model under the same endpoint familySource: Moonshot Kimi blog (kimi.com/blog/kimi-k2-6), HuggingFace moonshotai/Kimi-K2.6, Artificial Analysis, OpenRouter, SiliconANGLE · 2026-04-20
  • Self-hosting K2.5 / K2.6 at usable speed requires $30K+ in enterprise GPU hardware (8x H200 FP8 or 16x H100 production-grade) -- realistically this is a hosted-API model. Mac Studio M3 Ultra 256 GB unified RAM at Q2 quantization runs the model but at ~3 tok/sSource: Reddit r/LocalLLaMA, llm-stats.com · 2026-03
  • Early K2.5 releases had inconsistent tool-calling when quantized below Q4 -- community fixes landed March 2026; K2.6 inherits the same tool-use stack so quant guidance carries forwardSource: Hugging Face discussions · 2026-03

Best for

Agentic coding workflows, tool-use agents, long-horizon repository work (1M context), and teams who want frontier-tier quality at a fraction of frontier pricing ($15/M output vs Fable 5's $50/M).

Not for

Solo developers or hobbyists who want to run models locally -- the K3 weights are public now, but 2.8T parameters is datacenter territory, far beyond consumer hardware. Use Qwen3-Coder-Next or DeepSeek for self-hosting today.

Our Verdict

Kimi K3 (July 16-17, 2026) vaulted Moonshot from 'best open-weights value' to genuine frontier contention: 2.8T parameters, 1M context, multimodal, ranked best-available on Arena.AI at launch, and priced at $3/$15 per 1M -- a fifth of Anthropic's Fable 5 output rate. The launch rattled markets enough to be called a second DeepSeek shock. The open-weight promise has since been kept: the weights landed on Hugging Face around July 26-27 at 2.8T total / 104B activated, though under a bespoke Kimi K3 License rather than the Modified MIT of the K2 line, so read the terms before commercial use. Early August settled the distribution question too -- K3 is now selectable inside GitHub Copilot (hosted on Fireworks AI) and Perplexity (US-hosted), which lets Western teams use it without touching Moonshot's API. Remaining caveats: vendor benchmark claims still await third-party verification, the Moonshot API is visibly capacity-strained, and at 2.8T parameters self-hosting is datacenter work, not a laptop project. If you want maximum capability per dollar, K3 is the first model to test -- via Copilot or Perplexity if jurisdiction is a concern, via the Moonshot API if it isn't.

Sources

  • GitHub Changelog: Kimi K2.7 Code deprecated from Copilot 2026-10-02, K3 the suggested alternative (2026-09-03) (accessed 2026-09-05)
  • GitHub changelog: Kimi K3 is now available in GitHub Copilot (2026-08-06) (accessed 2026-08-06)
  • Perplexity changelog: Shared workspaces, Personal Computer for Windows, and Model Council -- includes US-hosted Kimi K3 (2026-08-04) (accessed 2026-08-06)
  • Hugging Face: moonshotai/Kimi-K3 model card (weights published 2026-07-27; 2.8T total / 104B activated; Kimi K3 License) (accessed 2026-07-29)
  • OpenRouter: Kimi K3 (pricing, context, capacity notes) (accessed 2026-07-18)
  • Fortune: Moonshot Kimi K3 rattles markets (2026-07-17) (accessed 2026-07-18)
  • Moonshot Kimi K2.6 blog (GA 2026-04-20) (accessed 2026-04-27)
  • HuggingFace moonshotai/Kimi-K2.6 (accessed 2026-04-27)
  • Artificial Analysis: Kimi K2.6 leading open weights (accessed 2026-04-27)
  • SiliconANGLE: Kimi K2.6 release (accessed 2026-04-27)
  • OpenRouter Kimi K2.6 pricing (accessed 2026-04-27)
  • llm-stats.com (accessed 2026-04-13)
  • Reddit r/singularity, r/LocalLLaMA (accessed 2026-04-13)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Kimi K3 (Moonshot)

Llama 4 (Meta) logo

Llama 4 (Meta)

Meta's open-weights family -- Scout (10M context), Maverick (multimodal 400B MoE). NOTE: Meta's frontier work moved to the proprietary Muse Spark line in April 2026; Llama remains downloadable and supported but is effectively in maintenance mode

B
7.9/10
Free tierFrom $0
Llama 4 Scout has a 10M token context wi...Llama 4 Maverick is natively multimodal ...
Updated 2026-06-09
Mistral AI logo

Mistral AI

European AI lab with open and commercial models -- Le Chat is now **Vibe** (May 28 2026): one agent across Work Mode + Code Mode with a VS Code extension and CLI, powered by Mistral Medium 3.5 (128B dense, 256k context, 77.6% SWE-Bench Verified). Newest release: **Shieldstral 1.0** (Aug 4 2026), a 3B Apache 2.0 multimodal safety classifier that runs on one 16GB GPU. Earlier 2026 line: Small 4 (119B MoE Apache 2.0), Voxtral TTS. **Mistral Medium 3.1 and Medium 3 were both retired 2026-08-31** -- Medium 3.5 is the vendor-listed replacement for each

B
7.5/10
Free tierFrom $0
Mistral Medium 3.5 (April 29 2026) is Mi...Vibe Remote Agents (also 4/29) lets you ...
Updated 2026-08-31
DeepSeek logo

DeepSeek

DeepSeek V4 shipped 2026-04-24: V4-Pro (1.6T/49B active MoE) + V4-Flash (284B/13B active), 1M native context, Hybrid Attention Architecture, open-source on HF. **V4-Pro reached GA 2026-08-13** (Terminal Bench 2.1 87.9, three thinking-effort levels, native Responses API). **Pricing changed 2026-08-16 and the new rates are now live:** peak/off-peak billing raised rates at every hour of the day -- V4-Pro went from $0.435/$0.87 to **$0.66/$1.98 off-peak and $1.32/$3.96 peak**. Despite the 'off-peak' framing there is no hour at which it is cheaper than before

A
8.0/10
Free tierFrom $0
Pricing is absurdly cheap compared to GP...DeepSeek-R1 reasoning model genuinely co...
Updated 2026-08-28
Gemma 4 (Google) logo

Gemma 4 (Google)

Google DeepMind's open-weights model family -- multimodal, 256K context, runs on edge devices

A
8.3/10
Free tierFrom $0
Apache 2.0 license -- truly permissive, ...Multimodal: handles text + image input (...
Updated 2026-04-19
DiffusionGemma (Google) logo

DiffusionGemma (Google)

Google DeepMind's experimental open-weights TEXT-DIFFUSION model (June 10, 2026) -- 26B MoE (3.8B active), Apache 2.0, generates 256-token blocks in parallel with bidirectional attention for up to 4x faster output (1,000+ tok/s on H100). Trades some quality vs Gemma 4 for raw speed

C
6.8/10
Free tierFrom $0
Text diffusion instead of autoregression...First open-weights text-diffusion model ...
Updated 2026-06-10
Qwen (Alibaba) logo

Qwen (Alibaba)

Alibaba's open-weights + API family -- Qwen3.8-Max flagship previewed at WAIC (Jul 19 2026: 2.4T sparse-MoE multimodal, closed preview, 'second only to Fable 5'), Qwen 3.7 Max GA (SWE-Bench Pro 60.6%, Terminal-Bench 69.7%, $2.50/$7.50 per 1M), Qwen3.6-27B dense Apache 2.0 (beats the 397B MoE on coding from one consumer GPU)

A
8.8/10
Free tierFrom $0
Qwen 3.6-Plus (launched Mar 30 2026) is ...Qwen3.5 Small (0.8B / 2B / 4B / 9B) is t...
Updated 2026-07-22
GLM / Z.ai (Zhipu AI) logo

GLM / Z.ai (Zhipu AI)

Zhipu AI's open-weights flagship -- GLM-5.2 (launched 2026-06-13) is a ~753B-parameter MoE with a 1M-token context and the new IndexShare sparse-attention architecture (~2.9x lower per-token FLOPs at 1M context), MIT licensed. Vendor benchmarks put SWE-Bench Pro at 62.1 (up from GLM-5.1's 58.4) and it tops the Artificial Analysis open-weights Intelligence Index; VentureBeat reports it beats GPT-5.5 on several long-horizon coding benchmarks at roughly 1/6 the cost. Drop-in for Claude Code / Cline / OpenCode. Still trained outside the Nvidia stack on Huawei Ascend silicon

A
8.0/10
Free tierFrom $0
GLM-5.1 (2026-04-07) topped SWE-Bench Pr...First frontier model trained entirely on...
Updated 2026-07-09
Nemotron (Nvidia) logo

Nemotron (Nvidia)

Nvidia's open-weights family -- hybrid Mamba-Transformer MoE architecture, optimized for efficient reasoning on Nvidia hardware. Nemotron 3 Ultra (550B total / 55B active) shipped 2026-06-04 as the family flagship, joining Super (120B/12B, March) and Nano

B
7.8/10
Free tierFrom $0
Hybrid Mamba-Transformer architecture dr...Nemotron 3 Super activates only 12B of 1...
Updated 2026-07-05
MiniMax M3 logo

MiniMax M3

MiniMax's coding/agent flagship -- M3 (June 1 2026): 1M-token context, MSA sparse attention (>15x decoding speedup at long context), SWE-Bench Pro 59.0%, Terminal-Bench 66.0%. OPEN WEIGHTS LIVE on HuggingFace since June 12 (~428B total / ~23B active, native multimodal, minimax-community license)

A
8.4/10
Free tierFrom $0
229B/10B-active MoE delivers Tier-1 agen...Sparse MoE design: ~10B active params du...
Updated 2026-08-03
Falcon (TII) logo

Falcon (TII)

UAE's Technology Innovation Institute open-weights family -- Falcon 3 optimized for efficient sub-10B deployment on consumer hardware

B
7.1/10
Free tierFrom $0
Apache 2.0 license -- fully permissive f...Sub-10B sizes run on any consumer GPU or...
Updated 2026-04-13
gpt-oss (OpenAI) logo

gpt-oss (OpenAI)

OpenAI's FIRST open-weight models -- gpt-oss-120b (single 80GB GPU, near parity with o4-mini on reasoning) and gpt-oss-20b (runs on 16GB edge devices). Apache 2.0. Launched 2025-08-05. gpt-oss-safeguard ships in 2026 as the safety-tuned variant

A
8.1/10
Free tierFrom $0
First-ever OpenAI open-weight release --...gpt-oss-120b approaches o4-mini on reaso...
Updated 2026-04-17
IBM Granite 4.0 logo

IBM Granite 4.0

IBM's enterprise-focused open-weight family -- Granite 4.0 hybrid Mamba-2 + transformer architecture (70-80% memory reduction vs pure transformer), 3B to 32B sizes, Apache 2.0. First open model family to secure ISO 42001 certification. Nano 350M runs on CPU with 8-16GB RAM. 3B Vision variant landed 2026-04-01

A
8.2/10
Free tierFrom $0
Hybrid Mamba-2 + transformer architectur...Granite 4.0 Nano (350M and 1.5B) is genu...
Updated 2026-04-17
Arcee Trinity-Large-Thinking logo

Arcee Trinity-Large-Thinking

Arcee AI's US-made open-weight frontier reasoning model -- launched 2026-04-01. 398B total params, ~13B active. Sparse MoE (256 experts, 4 active = 1.56% routing). Apache 2.0, trained from scratch. #2 on PinchBench trailing only Claude 3.5 Opus. ~96% cheaper than Opus-4.6 on agentic tasks

A
8.1/10
Free tierFrom $0
Rare US-made frontier-tier open-weight r...Trained from scratch (not a fine-tune) a...
Updated 2026-04-17
Olmo 3 (AI2) logo

Olmo 3 (AI2)

Allen Institute for AI's fully-open frontier reasoning models -- Olmo 3 family (2025-11-20) includes 7B and 32B sizes, four variants (Base, Think, Instruct, RLZero). Apache 2.0 with fully open data + checkpoints + training logs. Olmo 3-Think 32B matches Qwen3-32B-Thinking at 6x fewer training tokens

B
7.9/10
Free tierFrom $0
FULLY OPEN is a different category than ...Olmo 3-Think 32B matches Qwen3-32B-Think...
Updated 2026-04-17
AI21 Jamba2 logo

AI21 Jamba2

AI21 Labs' hybrid SSM-Transformer (Mamba-style) open-weight family -- Jamba2 launched 2026-01-08. Two sizes: 3B dense (runs on phones / laptops) and Jamba2 Mini MoE (12B active / 52B total). Apache 2.0, 256K context, mid-trained on 500B tokens

A
8.0/10
Free tierFrom $0
Hybrid SSM-Transformer (Mamba-style) arc...Jamba2 3B dense runs realistically on iP...
Updated 2026-04-17
StepFun Step 3.7 Flash logo

StepFun Step 3.7 Flash

StepFun's (China) agent-focused open-weight family -- Step 3.7 Flash (May 28 2026): 198B sparse MoE vision-language model, ~11B active, 256K context, Apache 2.0, ~400 tok/s, SWE-Bench Pro 56.3. Supersedes Step 3.5 Flash (Feb 2026) as the flagship

B
7.8/10
Free tierFrom $0
Step 3.5 Flash at 196B total / 11B activ...Agent-focused tuning explicitly -- tool ...
Updated 2026-06-10
Cohere Command A logo

Cohere Command A

Cohere's enterprise-multilingual flagship -- 111B params, 256K context, runs on 2x H100. 23 languages. CC-BY-NC 4.0 on weights (research / non-commercial), commercial requires Cohere enterprise contract. Follow-ups: Command A Reasoning + Command A Vision

B
7.5/10
Free tierFrom $0
Best-in-class multilingual open-weight m...Runs on just 2x H100 at FP16 for the ful...
Updated 2026-04-17
LongCat-2.0 (Meituan) logo

LongCat-2.0 (Meituan)

Meituan's open-source 1.6T-parameter MoE (~48B active) with native 1M-token context, MIT license -- trained entirely on domestic Chinese AI ASICs and revealed as the stealth 'Owl Alpha' model that had been topping OpenRouter

B
7.9/10
Free tierFrom $0
Genuinely open frontier scale: 1.6T tota...Native long context: LongCat Sparse Atte...
Updated 2026-07-05
Inkling (Thinking Machines Lab) logo

Inkling (Thinking Machines Lab)

Mira Murati's $12B lab ships its first model (2026-07-15): a 975B/41B-active open-weights MoE that reasons natively over text, images, and audio with a 1M-token context -- positioned not as the strongest model, but as the best starting point for fine-tuning via Tinker

A
8.0/10
Free tierFrom $0
First US frontier-adjacent open-weights ...Natively multimodal INPUT across text, i...
Updated 2026-07-18
Bonsai 27B (PrismML) logo

Bonsai 27B (PrismML)

The first 27B-class model that runs on a phone (2026-07-14) -- ternary and 1-bit quantizations of Qwen3.6 27B squeeze a multimodal, tool-calling, 262K-context model into 3.9-5.9GB under Apache 2.0

B
7.9/10
Free tierFrom $0
Genuine first: a 27B-class model running...Keeps the grown-up capabilities: multi-s...
Updated 2026-07-18