Kimi K3 (Moonshot) logo
A

Kimi K3 (Moonshot)

A Tier · 8.1/10

Moonshot's 2.8T-parameter Kimi K3 (launched 2026-07-16/17) is the largest open-weight model ever released -- 1M context, multimodal, $3/$15 per 1M via API, ranked best-available on Arena.AI at launch. WEIGHTS SHIPPED ~2026-07-26/27 on Hugging Face (2.8T total / 104B activated, safetensors) under a custom Kimi K3 License, not the Modified MIT of the K2 line

Last updated: 2026-08-06Free tier available

Score Breakdown

6.0
Ease of Use
9.0
Output Quality
8.5
Value
9.0
Features

Benchmark Scores

Benchmarks for Kimi K3 (2.8T, launched 2026-07-16) -- Arena.AI ranked it best-available at launch; vendor claims parity with Fable 5, third-party suites pending. Scores below are K2.6/K2.5-era baselines retained until K3 third-party runs publish

BenchmarkScore
SWE-Bench Pro58.6%
MMLU-Pro (K2.5 baseline)84.8%
GPQA Diamond (K2.5 baseline)80.5%
AIME 2025 (K2.5 baseline)91.2%
LiveCodeBench (K2.5 baseline)74.1%

Last updated: 2026-04-27

Personality & Tone

The long-context note-taker

Tone: Careful and document-focused. Kimi K2.5 shines when you dump a long document in -- replies read as summary-and-citation rather than open chat, leaning on the source material rather than the model's opinions.

Quirks: Context handling is the whole pitch. Without a document to anchor to, replies feel plainer than Qwen or DeepSeek. Native Chinese quality is very strong; English is decent but not class-leading.

The Good and the Bad

What we like

  • +Frontier-tier performance -- Elo 1309 on GDPval-AA, behind only OpenAI and Anthropic flagships
  • +Beats Claude Opus 4.5 on several coding benchmarks per community testing
  • +Unified thinking + non-thinking modes in one model (no need to swap)
  • +256K context window handles large codebases for agentic coding
  • +Weights are genuinely downloadable -- K2.6/K2.7-Code under Modified MIT, and K3 under its own Kimi K3 License (check its terms before commercial use)
  • +Native tool-use and agentic planning trained in -- not bolted on

What could be better

  • Self-hosting is datacenter-only -- K3 is 2.8T params (104B activated) and even the older K2 line needs 4+ H100-class GPUs
  • Moonshot is a smaller lab than DeepSeek/Alibaba -- less Western infrastructure support
  • API pricing ($0.60 in / $3.00 out) is higher than DeepSeek V3.2 ($0.28 in / $0.42 out)
  • PRC content filters apply (Tiananmen, Taiwan, etc.)
  • Documentation is heavily Chinese-first -- English docs trail releases

Pricing

API (Kimi K3)

$3 / $15/per 1M tokens (input/output)
  • K3 (launched 2026-07-16): 2.8T-param open-weight multimodal reasoning model
  • 1M token context window
  • Reasoning effort currently supports only 'max'
  • Live on kimi.com, Kimi app, Moonshot API, and OpenRouter; capacity-limited at launch (frequent 429s)

Self-hosted (Free -- K2.6/K2.7 line)

$0
  • K3 WEIGHTS PUBLISHED ~2026-07-26/27 at huggingface.co/moonshotai/Kimi-K3 -- safetensors, 2.8T total params, 104B activated, 1,048,576-token context
  • K3 ships under its own **Kimi K3 License**, NOT the Modified MIT used for K2 -- read the license terms before commercial use
  • K2.6 + K2.7-Code weights remain on Hugging Face under Modified MIT
  • Fine-tuning permitted

API (Moonshot direct, K2.6)

$0.60/per 1M input tokens
  • K2.6: $0.60 in / $2.50 out (Moonshot direct)
  • 256K context
  • Native video input (mp4/mov/avi/webm)

System Requirements

Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.

Model variantMinMax
Kimi K2.5 (1T total, 32B active MoE)Practically a hosted-only model for most users -- self-hosting requires enterprise hardware256 GB unified RAM Mac Studio M3 Ultra (Q2, ~3 tok/s)8× H200 141 GB FP8 or 16× H100 (production-grade)

Known Issues

  • K3 BREAKS INTO WESTERN PRODUCT DISTRIBUTION IN A SINGLE WEEK -- GITHUB COPILOT AND PERPLEXITY (2026-08-04 and 2026-08-06, both vendor-primary): three weeks after launch, K3 stopped being an API-and-weights story and became a model you can select inside mainstream Western products. **GITHUB COPILOT (2026-08-06):** K3 is **generally available** in the Copilot model picker across **Pro, Pro+, Max, Business and Enterprise** -- including base Pro, a broader entitlement than Claude Opus 5 got on 7/24. GitHub's own assessment: it '**shows frontier-level abilities on agentic coding with highly cost-effective pricing**.' Crucially, **GitHub hosts K3 on Fireworks AI**, so requests do not go to Moonshot's API. Billed at provider list pricing under usage-based billing; available in VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, JetBrains, Xcode, Eclipse and GitHub Mobile. It is **off by default for Business and Enterprise** and needs an explicit admin policy, and GitHub pointedly tells admins to 'review open-weight models against their own security, compliance, and data-governance requirements before enabling them.' **PERPLEXITY (2026-08-04):** K3 is available in Perplexity and Perplexity Computer for **Pro and Max** subscribers, '**hosted exclusively on US-based servers**' -- Perplexity's pitch is explicitly '2.8T-parameter mixture-of-experts model with a 1M-token context window and native vision, **without data leaving US infrastructure**.' **WHY THIS MATTERS MORE THAN A NORMAL INTEGRATION:** the single biggest practical objection to Kimi in Western enterprises was jurisdictional, not technical. Both integrations answer it the same way -- re-host the open weights on US infrastructure and sell the model without the Chinese API. That is only possible *because* the weights are public, which makes this the clearest commercial payoff yet for Moonshot's open-weight strategy. It does not, however, resolve the licence question: the **Kimi K3 License** still governs the weights themselves, so re-hosting by a vendor is not the same as your own right to deploy themSource: GitHub changelog (github.blog/changelog/2026-08-06-kimi-k3-is-now-available-in-github-copilot/, fetched 2026-08-06), Perplexity changelog (perplexity.ai/changelog/shared-workspaces-personal-computer-for-windows-and-model-council, dated 08/04/26, fetched 2026-08-06) · 2026-08-06
  • OPEN WEIGHTS SHIPPED (~2026-07-26/27 -- Moonshot announced 7/27 and reporting puts the actual drop a day earlier, so treat the exact day as approximate): Moonshot published **Kimi K3's weights** to Hugging Face at `moonshotai/Kimi-K3`, resolving the open question this page carried since launch. The model card confirms **2.8T total parameters with 104B activated** (the ~50B active figure that circulated in aggregator coverage was wrong), a **1,048,576-token context window**, and safetensors in F32 / BF16 / U8. **License is the bespoke 'Kimi K3 License', NOT the Modified MIT that covers K2.6 and K2.7-Code** -- the card states 'Both the code repository and the model weights are released under the Kimi K3 License', so anyone planning commercial use or a derivative needs to read those terms rather than assuming K2 permissions carry over. Practical effect: K3 becomes the largest open-weight model actually downloadable, though at 2.8T params it is a datacenter-class deployment, not a local oneSource: Hugging Face model card (huggingface.co/moonshotai/Kimi-K3, fetched 2026-07-29) · 2026-07-27
  • MODEL LAUNCH -- KIMI K3 (2026-07-16/17): Moonshot shipped **Kimi K3**, billed as the largest open-weight model ever announced -- **2.8T total parameters**, multimodal reasoning, **1M token context**, live immediately on kimi.com, the Kimi app, the Moonshot API, and OpenRouter at **$3/M input, $15/M output** (vs Fable 5's $50/M output). Unveiled at the World AI Conference in Shanghai 7/17 after appearing on platforms 7/16; the stealth Arena model 'Kivine' was K3 in testing, and Arena.AI ranked it the best available model at launch. Vendor claims it performs competitively with Claude Fable 5 and 'substantially outperforms' Opus 4.8 and GPT-5.6 Sol -- third-party verification pending; treat vendor benchmark claims accordingly. Market reaction was dubbed a 'second DeepSeek shock': TSMC fell 7%, SoftBank 9%, Nasdaq 100 ~1% on 7/17. CAVEATS (updated 2026-07-29 -- (a) and (d) are now RESOLVED, see the weights-release entry above): (a) weights shipped ~2026-07-26/27; (b) API capacity is limited at launch -- OpenRouter flags frequent 429 errors; (c) reasoning effort currently supports only 'max'; (d) the aggregator-reported ~50B active-parameter figure was WRONG -- the model card publishes **104B activated**Source: OpenRouter (openrouter.ai/moonshotai/kimi-k3), Fortune (2026-07-17), Reuters, CNBC, r/LocalLLaMA · 2026-07-17
  • MODEL LAUNCH (2026-06-12): **Kimi K2.7-Code** -- Moonshot's code-specialized variant dropped on HuggingFace (moonshotai/Kimi-K2.7-Code, HN front page 333 points). Specs: 1T-param MoE with 32B active (384 experts), 256K context, Modified MIT license, MoonViT 400M vision encoder, built on K2.6, ~30% fewer thinking tokens than K2.6, forces thinking mode on. API via platform.moonshot.ai (OpenAI- and Anthropic-compatible endpoints). Notably honest self-published benchmarks show it TRAILING the frontier: Kimi Code Bench v2 62.0 vs GPT-5.5's 69.0 and Opus 4.8's 67.4 -- Moonshot is positioning on open-weights value, not SOTA claims. API pricing circulating in aggregators (~$0.19 cached/$0.95 in/$4.00 out per 1M) -- verify on the vendor pricing page before relying on itSource: HuggingFace (huggingface.co/moonshotai/Kimi-K2.7-Code), Hacker News, platform.moonshot.ai · 2026-06-12
  • API DEPRECATION (2026-05-25, vendor docs verbatim: 'The kimi-k2 series models were officially discontinued on May 25, 2026'): retired model ids -- kimi-k2-0905-preview, kimi-k2-0711-preview, kimi-k2-turbo-preview, kimi-k2-thinking, kimi-k2-thinking-turbo. If your code pins any of these, requests now fail; migrate to kimi-k2.6 (or kimi-k2.5, which remains an active model alongside it). K2.6 detail confirmed on the vendor blog: open weights on HuggingFace (moonshotai/Kimi-K2.6), 256K context (262,144 default), agent swarm scaling to **300 sub-agents / 4,000 coordinated steps** (up from K2.5's 100/1,500)Source: Moonshot platform docs (platform.kimi.ai/docs/models), kimi.com/blog/kimi-k2-6, HuggingFace moonshotai/Kimi-K2.6 · 2026-05-25
  • SUPERSEDED (2026-07-16): the June-era 'K3 never shipped / treat K3 claims as fabrication' guidance no longer holds -- Kimi K3 is real and launched July 16-17, 2026 (see the K3 launch entry above). The May Manifold window did resolve NO (K3 missed May by six weeks), and the skepticism was correct at the time; the launch simply came later than the rumor mill claimed. Retained for history: the K2-series API deprecation (5/25) and K2.6 investment preceded K3 rather than replacing it.Source: kimi.com, OpenRouter (openrouter.ai/moonshotai/kimi-k3), Fortune (2026-07-17) · 2026-07-16
  • Kimi K2.6 (GA 2026-04-20) supersedes K2.5 -- 1T total / 32B active MoE, 256K context, adds native video input (mp4/mov/avi/webm). Scores 54 on Artificial Analysis Intelligence Index v4.0, ranked #1 open-weights and #4 overall (three points behind Claude Opus 4.7 / Gemini 3.1 Pro / OpenAI flagships at 57). SWE-Bench Pro 58.6%. Modified MIT license unchanged. Moonshot direct API: $0.60 in / $2.50 out per 1M tokens. OpenRouter blended: ~$0.95 in / $4.00 out. If you were on K2.5, the upgrade is non-breaking on the API side -- Moonshot routes the K2.6 model under the same endpoint familySource: Moonshot Kimi blog (kimi.com/blog/kimi-k2-6), HuggingFace moonshotai/Kimi-K2.6, Artificial Analysis, OpenRouter, SiliconANGLE · 2026-04-20
  • Self-hosting K2.5 / K2.6 at usable speed requires $30K+ in enterprise GPU hardware (8x H200 FP8 or 16x H100 production-grade) -- realistically this is a hosted-API model. Mac Studio M3 Ultra 256 GB unified RAM at Q2 quantization runs the model but at ~3 tok/sSource: Reddit r/LocalLLaMA, llm-stats.com · 2026-03
  • Early K2.5 releases had inconsistent tool-calling when quantized below Q4 -- community fixes landed March 2026; K2.6 inherits the same tool-use stack so quant guidance carries forwardSource: Hugging Face discussions · 2026-03

Best for

Agentic coding workflows, tool-use agents, long-horizon repository work (1M context), and teams who want frontier-tier quality at a fraction of frontier pricing ($15/M output vs Fable 5's $50/M).

Not for

Solo developers or hobbyists who want to run models locally -- the K3 weights are public now, but 2.8T parameters is datacenter territory, far beyond consumer hardware. Use Qwen3-Coder-Next or DeepSeek for self-hosting today.

Our Verdict

Kimi K3 (July 16-17, 2026) vaulted Moonshot from 'best open-weights value' to genuine frontier contention: 2.8T parameters, 1M context, multimodal, ranked best-available on Arena.AI at launch, and priced at $3/$15 per 1M -- a fifth of Anthropic's Fable 5 output rate. The launch rattled markets enough to be called a second DeepSeek shock. The open-weight promise has since been kept: the weights landed on Hugging Face around July 26-27 at 2.8T total / 104B activated, though under a bespoke Kimi K3 License rather than the Modified MIT of the K2 line, so read the terms before commercial use. Early August settled the distribution question too -- K3 is now selectable inside GitHub Copilot (hosted on Fireworks AI) and Perplexity (US-hosted), which lets Western teams use it without touching Moonshot's API. Remaining caveats: vendor benchmark claims still await third-party verification, the Moonshot API is visibly capacity-strained, and at 2.8T parameters self-hosting is datacenter work, not a laptop project. If you want maximum capability per dollar, K3 is the first model to test -- via Copilot or Perplexity if jurisdiction is a concern, via the Moonshot API if it isn't.

Sources

  • GitHub changelog: Kimi K3 is now available in GitHub Copilot (2026-08-06) (accessed 2026-08-06)
  • Perplexity changelog: Shared workspaces, Personal Computer for Windows, and Model Council -- includes US-hosted Kimi K3 (2026-08-04) (accessed 2026-08-06)
  • Hugging Face: moonshotai/Kimi-K3 model card (weights published 2026-07-27; 2.8T total / 104B activated; Kimi K3 License) (accessed 2026-07-29)
  • OpenRouter: Kimi K3 (pricing, context, capacity notes) (accessed 2026-07-18)
  • Fortune: Moonshot Kimi K3 rattles markets (2026-07-17) (accessed 2026-07-18)
  • Moonshot Kimi K2.6 blog (GA 2026-04-20) (accessed 2026-04-27)
  • HuggingFace moonshotai/Kimi-K2.6 (accessed 2026-04-27)
  • Artificial Analysis: Kimi K2.6 leading open weights (accessed 2026-04-27)
  • SiliconANGLE: Kimi K2.6 release (accessed 2026-04-27)
  • OpenRouter Kimi K2.6 pricing (accessed 2026-04-27)
  • llm-stats.com (accessed 2026-04-13)
  • Reddit r/singularity, r/LocalLLaMA (accessed 2026-04-13)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Kimi K3 (Moonshot)

Llama 4 (Meta) logo

Llama 4 (Meta)

Meta's open-weights family -- Scout (10M context), Maverick (multimodal 400B MoE). NOTE: Meta's frontier work moved to the proprietary Muse Spark line in April 2026; Llama remains downloadable and supported but is effectively in maintenance mode

B
7.9/10
Free tierFrom $0
Llama 4 Scout has a 10M token context wi...Llama 4 Maverick is natively multimodal ...
Updated 2026-06-09
Mistral AI logo

Mistral AI

European AI lab with open and commercial models -- Le Chat is now **Vibe** (May 28 2026): one agent across Work Mode + Code Mode with a VS Code extension and CLI, powered by Mistral Medium 3.5 (128B dense, 256k context, 77.6% SWE-Bench Verified). Newest release: **Shieldstral 1.0** (Aug 4 2026), a 3B Apache 2.0 multimodal safety classifier that runs on one 16GB GPU. Earlier 2026 line: Small 4 (119B MoE Apache 2.0), Voxtral TTS. **Mistral Medium 3.1 and Medium 3 were both retired 2026-08-31** -- Medium 3.5 is the vendor-listed replacement for each

B
7.5/10
Free tierFrom $0
Mistral Medium 3.5 (April 29 2026) is Mi...Vibe Remote Agents (also 4/29) lets you ...
Updated 2026-08-31
DeepSeek logo

DeepSeek

DeepSeek V4 shipped 2026-04-24: V4-Pro (1.6T/49B active MoE) + V4-Flash (284B/13B active), 1M native context, Hybrid Attention Architecture, open-source on HF. **V4-Pro reached GA 2026-08-13** (Terminal Bench 2.1 87.9, three thinking-effort levels, native Responses API). **Pricing changed 2026-08-16 and the new rates are now live:** peak/off-peak billing raised rates at every hour of the day -- V4-Pro went from $0.435/$0.87 to **$0.66/$1.98 off-peak and $1.32/$3.96 peak**. Despite the 'off-peak' framing there is no hour at which it is cheaper than before

A
8.0/10
Free tierFrom $0
Pricing is absurdly cheap compared to GP...DeepSeek-R1 reasoning model genuinely co...
Updated 2026-08-28
Gemma 4 (Google) logo

Gemma 4 (Google)

Google DeepMind's open-weights model family -- multimodal, 256K context, runs on edge devices

A
8.3/10
Free tierFrom $0
Apache 2.0 license -- truly permissive, ...Multimodal: handles text + image input (...
Updated 2026-04-19
DiffusionGemma (Google) logo

DiffusionGemma (Google)

Google DeepMind's experimental open-weights TEXT-DIFFUSION model (June 10, 2026) -- 26B MoE (3.8B active), Apache 2.0, generates 256-token blocks in parallel with bidirectional attention for up to 4x faster output (1,000+ tok/s on H100). Trades some quality vs Gemma 4 for raw speed

C
6.8/10
Free tierFrom $0
Text diffusion instead of autoregression...First open-weights text-diffusion model ...
Updated 2026-06-10
Qwen (Alibaba) logo

Qwen (Alibaba)

Alibaba's open-weights + API family -- Qwen3.8-Max flagship previewed at WAIC (Jul 19 2026: 2.4T sparse-MoE multimodal, closed preview, 'second only to Fable 5'), Qwen 3.7 Max GA (SWE-Bench Pro 60.6%, Terminal-Bench 69.7%, $2.50/$7.50 per 1M), Qwen3.6-27B dense Apache 2.0 (beats the 397B MoE on coding from one consumer GPU)

A
8.8/10
Free tierFrom $0
Qwen 3.6-Plus (launched Mar 30 2026) is ...Qwen3.5 Small (0.8B / 2B / 4B / 9B) is t...
Updated 2026-07-22
GLM / Z.ai (Zhipu AI) logo

GLM / Z.ai (Zhipu AI)

Zhipu AI's open-weights flagship -- GLM-5.2 (launched 2026-06-13) is a ~753B-parameter MoE with a 1M-token context and the new IndexShare sparse-attention architecture (~2.9x lower per-token FLOPs at 1M context), MIT licensed. Vendor benchmarks put SWE-Bench Pro at 62.1 (up from GLM-5.1's 58.4) and it tops the Artificial Analysis open-weights Intelligence Index; VentureBeat reports it beats GPT-5.5 on several long-horizon coding benchmarks at roughly 1/6 the cost. Drop-in for Claude Code / Cline / OpenCode. Still trained outside the Nvidia stack on Huawei Ascend silicon

A
8.0/10
Free tierFrom $0
GLM-5.1 (2026-04-07) topped SWE-Bench Pr...First frontier model trained entirely on...
Updated 2026-07-09
Nemotron (Nvidia) logo

Nemotron (Nvidia)

Nvidia's open-weights family -- hybrid Mamba-Transformer MoE architecture, optimized for efficient reasoning on Nvidia hardware. Nemotron 3 Ultra (550B total / 55B active) shipped 2026-06-04 as the family flagship, joining Super (120B/12B, March) and Nano

B
7.8/10
Free tierFrom $0
Hybrid Mamba-Transformer architecture dr...Nemotron 3 Super activates only 12B of 1...
Updated 2026-07-05
MiniMax M3 logo

MiniMax M3

MiniMax's coding/agent flagship -- M3 (June 1 2026): 1M-token context, MSA sparse attention (>15x decoding speedup at long context), SWE-Bench Pro 59.0%, Terminal-Bench 66.0%. OPEN WEIGHTS LIVE on HuggingFace since June 12 (~428B total / ~23B active, native multimodal, minimax-community license)

A
8.4/10
Free tierFrom $0
229B/10B-active MoE delivers Tier-1 agen...Sparse MoE design: ~10B active params du...
Updated 2026-08-03
Falcon (TII) logo

Falcon (TII)

UAE's Technology Innovation Institute open-weights family -- Falcon 3 optimized for efficient sub-10B deployment on consumer hardware

B
7.1/10
Free tierFrom $0
Apache 2.0 license -- fully permissive f...Sub-10B sizes run on any consumer GPU or...
Updated 2026-04-13
gpt-oss (OpenAI) logo

gpt-oss (OpenAI)

OpenAI's FIRST open-weight models -- gpt-oss-120b (single 80GB GPU, near parity with o4-mini on reasoning) and gpt-oss-20b (runs on 16GB edge devices). Apache 2.0. Launched 2025-08-05. gpt-oss-safeguard ships in 2026 as the safety-tuned variant

A
8.1/10
Free tierFrom $0
First-ever OpenAI open-weight release --...gpt-oss-120b approaches o4-mini on reaso...
Updated 2026-04-17
IBM Granite 4.0 logo

IBM Granite 4.0

IBM's enterprise-focused open-weight family -- Granite 4.0 hybrid Mamba-2 + transformer architecture (70-80% memory reduction vs pure transformer), 3B to 32B sizes, Apache 2.0. First open model family to secure ISO 42001 certification. Nano 350M runs on CPU with 8-16GB RAM. 3B Vision variant landed 2026-04-01

A
8.2/10
Free tierFrom $0
Hybrid Mamba-2 + transformer architectur...Granite 4.0 Nano (350M and 1.5B) is genu...
Updated 2026-04-17
Arcee Trinity-Large-Thinking logo

Arcee Trinity-Large-Thinking

Arcee AI's US-made open-weight frontier reasoning model -- launched 2026-04-01. 398B total params, ~13B active. Sparse MoE (256 experts, 4 active = 1.56% routing). Apache 2.0, trained from scratch. #2 on PinchBench trailing only Claude 3.5 Opus. ~96% cheaper than Opus-4.6 on agentic tasks

A
8.1/10
Free tierFrom $0
Rare US-made frontier-tier open-weight r...Trained from scratch (not a fine-tune) a...
Updated 2026-04-17
Olmo 3 (AI2) logo

Olmo 3 (AI2)

Allen Institute for AI's fully-open frontier reasoning models -- Olmo 3 family (2025-11-20) includes 7B and 32B sizes, four variants (Base, Think, Instruct, RLZero). Apache 2.0 with fully open data + checkpoints + training logs. Olmo 3-Think 32B matches Qwen3-32B-Thinking at 6x fewer training tokens

B
7.9/10
Free tierFrom $0
FULLY OPEN is a different category than ...Olmo 3-Think 32B matches Qwen3-32B-Think...
Updated 2026-04-17
AI21 Jamba2 logo

AI21 Jamba2

AI21 Labs' hybrid SSM-Transformer (Mamba-style) open-weight family -- Jamba2 launched 2026-01-08. Two sizes: 3B dense (runs on phones / laptops) and Jamba2 Mini MoE (12B active / 52B total). Apache 2.0, 256K context, mid-trained on 500B tokens

A
8.0/10
Free tierFrom $0
Hybrid SSM-Transformer (Mamba-style) arc...Jamba2 3B dense runs realistically on iP...
Updated 2026-04-17
StepFun Step 3.7 Flash logo

StepFun Step 3.7 Flash

StepFun's (China) agent-focused open-weight family -- Step 3.7 Flash (May 28 2026): 198B sparse MoE vision-language model, ~11B active, 256K context, Apache 2.0, ~400 tok/s, SWE-Bench Pro 56.3. Supersedes Step 3.5 Flash (Feb 2026) as the flagship

B
7.8/10
Free tierFrom $0
Step 3.5 Flash at 196B total / 11B activ...Agent-focused tuning explicitly -- tool ...
Updated 2026-06-10
Cohere Command A logo

Cohere Command A

Cohere's enterprise-multilingual flagship -- 111B params, 256K context, runs on 2x H100. 23 languages. CC-BY-NC 4.0 on weights (research / non-commercial), commercial requires Cohere enterprise contract. Follow-ups: Command A Reasoning + Command A Vision

B
7.5/10
Free tierFrom $0
Best-in-class multilingual open-weight m...Runs on just 2x H100 at FP16 for the ful...
Updated 2026-04-17
LongCat-2.0 (Meituan) logo

LongCat-2.0 (Meituan)

Meituan's open-source 1.6T-parameter MoE (~48B active) with native 1M-token context, MIT license -- trained entirely on domestic Chinese AI ASICs and revealed as the stealth 'Owl Alpha' model that had been topping OpenRouter

B
7.9/10
Free tierFrom $0
Genuinely open frontier scale: 1.6T tota...Native long context: LongCat Sparse Atte...
Updated 2026-07-05
Inkling (Thinking Machines Lab) logo

Inkling (Thinking Machines Lab)

Mira Murati's $12B lab ships its first model (2026-07-15): a 975B/41B-active open-weights MoE that reasons natively over text, images, and audio with a 1M-token context -- positioned not as the strongest model, but as the best starting point for fine-tuning via Tinker

A
8.0/10
Free tierFrom $0
First US frontier-adjacent open-weights ...Natively multimodal INPUT across text, i...
Updated 2026-07-18
Bonsai 27B (PrismML) logo

Bonsai 27B (PrismML)

The first 27B-class model that runs on a phone (2026-07-14) -- ternary and 1-bit quantizations of Qwen3.6 27B squeeze a multimodal, tool-calling, 262K-context model into 3.9-5.9GB under Apache 2.0

B
7.9/10
Free tierFrom $0
Genuine first: a 27B-class model running...Keeps the grown-up capabilities: multi-s...
Updated 2026-07-18