Hunyuan 3 (Tencent Hy3) logo
A

Hunyuan 3 (Tencent Hy3)

A Tier · 8.1/10

Tencent's Hy3 reached GA 2026-07-06 (upgraded from the April preview) -- 295B total / 21B active MoE, 256K context, now Apache 2.0 open weights on HuggingFace + ModelScope with the EU/UK/South Korea restriction lifted. ~90% agent-task completion on Tencent's internal apps; API via Tencent Cloud TokenHub. Integrated into Yuanbao, WeChat, QQ

Last updated: 2026-07-22Free tier available

Score Breakdown

7.0
Ease of Use
8.0
Output Quality
9.5
Value
8.0
Features

Personality & Tone

Tencent's first serious open-weights LLM

Tone: Direct and structured, similar to DeepSeek's chat tone -- compact answers in technical domains, reasonable Chinese-language polish, English noticeably less expressive than Claude or GPT.

Quirks: PRC content filters apply across the same set of regulated topics as DeepSeek and Qwen. Distribution via Yuanbao + WeChat + QQ means many Hy3 users in the wild are talking to it through chat-app surfaces rather than a developer console.

The Good and the Bad

What we like

  • +Open weights from a top-3 Chinese tech company is itself the headline -- Tencent had been a notable holdout while DeepSeek, Alibaba, Moonshot, MiniMap, and Z.ai shipped open-weight flagships through 2024-25
  • +Pricing is aggressive. ~1.2 RMB per million input tokens puts Hy3 in the same value bracket as DeepSeek V4-Flash and well below Western frontier models
  • +256K context is comfortable for retrieval-heavy and agentic workflows -- not as long as the 1M-context cohort (DeepSeek V4, Qwen 3.6-Plus, Gemini 2.5) but ample for most real workloads
  • +Distribution moat: Tencent already has Yuanbao + WeChat + QQ surfaces with hundreds of millions of MAUs. The model meets users where they already are, which matters in the Chinese market more than benchmarks
  • +Hy3 is the start of a multi-model family per Tencent's launch comms -- expect Pro / Flash / multimodal variants in coming months along the DeepSeek and Alibaba pattern

What could be better

  • Third-party benchmarks (Artificial Analysis, LMSYS, SWE-bench) are still spotty as of launch week -- treat Tencent's own benchmark claims with the usual self-reporting discount until independent runs corroborate
  • PRC content filters apply -- expect refusals or blandness on Tiananmen, Taiwan sovereignty, Xi Jinping, and other politically sensitive topics. Pattern matches DeepSeek and Qwen
  • Geo-availability improved at GA: the preview license's explicit EU / UK / South Korea exclusion was lifted with the Apache 2.0 GA release (2026-07-06). Developers should still check local data-residency rules before deploying via API; self-hosted open-weights deployment remains the fallback
  • English-language polish lags Claude and GPT for nuanced writing tasks -- this is a coding / reasoning / Chinese-language pick first, prose-quality pick second

Pricing

Free (consumer)

$0
  • Yuanbao app (Tencent's consumer chatbot)
  • Inside WeChat + QQ via Tencent's AI integrations
  • Web access via hy3ai.com
  • Basic usage limits apply

API (Hy3 Preview)

~1.2 RMB/per 1M input tokens
  • Roughly $0.16 USD/M input at April 2026 exchange rates -- among the cheapest frontier-class APIs on the market
  • 295B total / 21B active MoE
  • 256K context window
  • Pricing position is closer to DeepSeek V4-Flash than to GPT-5.5 or Claude Opus 4.7

Self-hosted (open-source)

$0 + GPU costs
  • Open weights on Hugging Face (tencent/Hy3-preview)
  • Permits self-hosting and fine-tuning per the published license
  • 21B-active MoE means inference is feasible on multi-GPU consumer-tier setups for testing; production deployments still want H100/H200

System Requirements

Hardware needed to self-host. Min = smallest viable setup (usually heavy quantization). Max = full-precision / production-grade.

Model variantMinMax
Hy3 Preview (295B total, 21B active MoE)Open weights on HuggingFace (tencent/Hy3-preview). 21B-active means inference is cheaper than the 295B total suggests; multi-GPU desktop setups are feasible for evaluation, production wants datacenter GPUs128 GB RAM + 1× RTX 4090 (severe quantization, experimental)4× H100 FP8 or 2× H200 (production)

Known Issues

  • HY3 REACHED GA (2026-07-06): Tencent officially released the final Hunyuan Hy3 model, upgrading the April preview with more post-training compute and higher-quality data. Same architecture (295B total / 21B active MoE, 256K context) but two material changes: (1) it now ships under **Apache 2.0** with the preview's **EU / UK / South Korea use restriction lifted**, open-weights day-one on Hugging Face + ModelScope; (2) heavily strengthened autonomous-agent capability -- Tencent claims a **~90% task-completion rate** across several of its internal applications, and bundled a free AI-agent feature into the consumer apps. API is now served via **Tencent Cloud's TokenHub** platform. This flips the earlier 'preview-stage, not GA' caveat -- Hy3 is now a stable, production-oriented releaseSource: Tencent (tencent.com/en-us/articles/2202386.html), Caixin Global (2026-07-06), TechNode (2026-07-07), Pandaily · 2026-07-06
  • Pricing in RMB is the published number -- USD pricing depends on FX and may move. ~1.2 RMB/M input ≈ $0.16 USD at April 2026 rates. Output token pricing is published separately; check api docs before forecasting costSource: Caixin Global, Tencent product page · 2026-04
  • PRC content filtering applies. Hy3 declines to engage on Tiananmen Square, Taiwan sovereignty, and other regulated topics. Same operating reality as DeepSeek, Qwen, Kimi, GLM -- if your use case needs unfiltered geopolitics, choose a Western frontier model or a community-tuned open-weights variantSource: Pattern from DeepSeek/Qwen; Hy3 follows the same regulatory regime · 2026-04

Best for

Chinese-market builders, multilingual products that need strong CN performance, and cost-sensitive developers who want frontier-class quality at sub-DeepSeek pricing. Also anyone watching the open-weights field who wants Tencent's distribution-backed entry on the bench.

Not for

English-first creative writing (Claude or GPT-5.5 are stronger), use cases that touch Chinese-government-sensitive topics, or teams that need a fully proven track record of third-party benchmark verification. As preview-stage software it is genuinely 'preview.'

Our Verdict

Hy3 is Tencent finally entering the open-weights race with a serious model and aggressive pricing. The 295B / 21B-active MoE design is aligned with what Alibaba and DeepSeek have shown works at scale, and the ~1.2 RMB/M-token list price puts Hy3 in real competition with DeepSeek V4-Flash for the cost-per-token crown. Distribution through Yuanbao + WeChat + QQ is a structural advantage Western developers tend to under-weight. Caveats: it's preview-stage, third-party benchmarks are still incoming, PRC content filters apply. For a US/EU developer comparison-shopping a cheap frontier-class API right now, DeepSeek V4-Flash is still the safer first stop -- but Hy3 deserves a spot on the bench, especially if Chinese market reach matters.

Sources

  • Tencent: Hunyuan officially releases Hy3 (GA, 2026-07-06) (accessed 2026-07-22)
  • Caixin Global: Tencent launches final Hunyuan 3 model with free AI-agent feature (2026-07-06) (accessed 2026-07-22)
  • TechNode: Tencent launches Hunyuan Hy3, integrates across products (2026-07-07) (accessed 2026-07-22)
  • Caixin Global: Tencent unveils new AI model to close gap with rivals (2026-04-23) (accessed 2026-04-25)
  • HuggingFace: tencent/Hy3-preview (accessed 2026-04-25)
  • Pandaily: How Tencent is building its global AI moat (accessed 2026-04-25)
  • Dataconomy: Tencent uses product rollout to define Hy3 Preview (accessed 2026-04-25)
  • Hy3 product site (accessed 2026-04-25)

The Tier List Tuesday

Weekly newsletter: tier movers, new entrants, and the VS of the week. Built from our daily AI-tool sweeps. No spam, unsubscribe anytime.

Alternatives to Hunyuan 3 (Tencent Hy3)

Claude (Anthropic) logo

Claude (Anthropic)

**Life Sciences Verification Program opened 2026-09-17** -- verified life-science teams get Mythos 5.1, Opus 5 and Sonnet 5 with the biology safeguards relaxed (Standard Use grants renewed yearly; High-risk Use grants per project every six months, Opus 5 and Sonnet 5 today, Mythos still US-government-gated), on top of the 2026-08-27 Claude team plan for scientists (10,000 seats, standard free, premium $15/mo). Anthropic's flagship LLM family. **Claude Opus 5 launched 2026-07-24** and is now the default model on Claude Max and the strongest model on Claude Pro -- same $5/$25 per 1M as Opus 4.8, but Anthropic says it lands within 0.5% of Fable 5 on CursorBench at half the cost. **Sonnet 5's $2/$10 per 1M is now permanent** -- Anthropic cancelled the 2026-09-01 rise to $3/$15 and made the launch rate standard -- and it stays the default on Free/Pro. **Claude Fable 5.1 and Mythos 5.1 launched 2026-09-01** and now top the range -- same $10/$50 per 1M as Fable 5, with the saving delivered entirely through a 4x cheaper cache read ($1 -> $0.25/MTok), so it is 25-45% cheaper only if your workload reuses cached context. **From 2026-08-14 future Claude models watermark their text output globally** (SynthID-Text; no extra tokens, no price or speed change, no identifying information -- detector API not shipped yet), and the **legacy Workbench plus the experimental prompt-tools APIs retired 2026-08-17**

A
8.5/10
Free tierFrom $0
Best writing quality of any LLM -- Opus ...1M token context window for enterprise A...
Updated 2026-09-17
Claude Mythos 5.1 logo

Claude Mythos 5.1

Anthropic's trusted-access frontier model. **Mythos 5.1 launched 2026-09-01** alongside Fable 5.1, and Anthropic now states outright that they are **the same model with different safeguards** -- the 5.1-cycle gap is 60.9% vs 55.8% on Terminal-Bench 4.0, which Anthropic attributes to safeguard interventions rather than capability and expects to shrink. Originally launched June 9, 2026 alongside Claude Fable 5. Suspended June 12 by a US export-control order, then PARTIALLY RESTORED July 1, 2026 (US government lifted controls June 30): Mythos 5 is back for a set of US organizations with government approval, while Anthropic works to re-expand the broader Glasswing program. Public Fable 5 returned globally the same day. Gated to Project Glasswing orgs + select biology researchers.

C
6.5/10
From Invite only
The most capable Anthropic model availab...73% success rate on expert-level Capture...
Updated 2026-09-05
Gemini (Google) logo

Gemini (Google)

**Gemini 3.8 Live and 3.8 Live Extended Thinking launched 2026-09-15** -- native speech-to-speech models in the Gemini API at $0.005/min audio in and $0.018/min audio out (audio-in a tenth of GPT-Live-1's $0.05/min voice layer, and on the same price row as the older 3.1 Flash Live), with Extended Thinking taking the #1 spot on Artificial Analysis' Speech to Speech Quality Index (82.6) and rolling into Gemini Live, Search Live and Workspace Docs/Gmail/Keep. **Gemini app for Windows shipped 2026-09-10** (Alt+Space overlay, Google-app connectors, Gemini Spark on desktop) and the Sept 9 AI-plan refresh added Google Pics and Sheets canvas for AI Pro/Ultra plus Gemini Spark hooks into Chrome and Google Photos. Google's LLM with deep Google Workspace integration, 2M token context window, and native code execution -- **Gemini 3.8 Flash launched 2026-09-02** -- the third Flash release in six weeks -- at the SAME $0.75/$3.75 per 1M intro rate as 3.7 Flash and, critically, **on the same 2026-12-31 expiry** (then $1.50/$7.50), so the discount window did not reset. HLE-Verified 54.9%; Google warns it uses more tokens per task because it 'works harder', so cost per task can rise at an unchanged rate. **Gemini 3.8 Flash Cyber** is trusted-defenders-only via the new **Fairwind Program**. Gemini 3.7 Flash (2026-08-13) remains fully supported for efficiency-first work. Gemini 3.5 Pro STILL delayed and partner-testing-only (Bloomberg, 7/16 -- coding shortfalls, no ship date), so Flash keeps shipping while the Pro line stalls. Gemini 4 pre-training underway. **Two new models in Aug 2026: Gemini 3.5 Transcribe (2026-08-26) speech-to-text and Gemini Omni 1.1 Flash (2026-08-27) production video with 40s scene extension, start/end frames and 4K** -- Transcribe's rate card has since appeared (about $0.005/min blended, verified 9/17); Omni 1.1 Flash's still has not

A
8.3/10
Free tierFrom $0
2 million token context window is the la...Best Google Workspace integration (Gmail...
Updated 2026-09-17
Grok logo

Grok

SpaceXAI's irreverent chatbot with a direct line to X/Twitter -- and now **Grok 4.6 (launched 2026-08-12)**, focused on long-running agents and interactive/visual work, at the same $2/$6 per 1M as Grok 4.5 (fast variant 2x). xAI says it matches GPT-5.6 Sol on the AA Intelligence Index (61), though Sol still leads it on DeepSWE and Terminal-Bench. **Grok Bot** (8/11, early beta) adds always-on agents with their own cloud computer. **Grok 4.6 reached GitHub Copilot on 2026-08-14** across all five paid Copilot SKUs (off by default for orgs), adding to its day-one Cursor availability. Grok 4.3 remains the value tier at $1.25/$2.50

B
7.5/10
Free tierFrom $0
Real-time access to X/Twitter data is ge...Grok 3 benchmarks are competitive with G...
Updated 2026-09-05
Muse Spark (Meta) logo

Muse Spark (Meta)

Meta's frontier model line from its Superintelligence Lab -- Muse Spark 1.1 (2026-07-09) adds substantially better coding, 1M-token multi-agent orchestration, and Meta's first paid developer API (Meta Model API, public preview)

A
8.8/10
Free tierFrom $0
Completely free to use via Meta AI app a...Natively multimodal: handles text, image...
Updated 2026-07-18
GPT-Rosalind (OpenAI) logo

GPT-Rosalind (OpenAI)

OpenAI's first domain-specific model -- life sciences, drug discovery, translational medicine. Launched 2026-04-16 as a Trusted Access research preview. Launch partners: Amgen, Moderna, Allen Institute, Thermo Fisher. Paired with a Life Sciences Codex plugin (50+ scientific tool integrations)

C
6.8/10
From Invite only
OpenAI's first named vertical/domain-spe...Launch partners Amgen, Moderna, Allen In...
Updated 2026-04-17
GPT-5.6-Cyber / GPT-5.4-Cyber (OpenAI) logo

GPT-5.6-Cyber / GPT-5.4-Cyber (OpenAI)

**OpenAI is retiring GPT-5.4-Cyber (notice 2026-09-11, removed from the API 2026-10-01); its replacement is `gpt-5.6-cyber`**, an alias for OpenAI's most advanced purpose-trained cyber models, gated behind the Daybreak program and priced at $12.50/$75 per 1M (first published price for any OpenAI cyber model). Original page: OpenAI's defensive-cybersecurity variant of GPT-5.4, launched 2026-04-16. Lowered refusal boundary for security-research tasks and native binary reverse-engineering. Access gated via Trusted Access for Cyber (TAC) program -- thousands of verified defenders, hundreds of teams, no public pricing. On **2026-08-17 OpenAI published its first dedicated post on the Hugging Face incident**, conceding it 'underestimated the real-world cyber capabilities of our AI models' and confirming it now releases cyber capabilities only to trusted defenders. **On 2026-09-01 OpenAI confirmed GPT-6 Astra meets the Critical cyber threshold -- the first model it has ever designated at that level -- and on 2026-09-03 committed $1B to Daybreak for Frontline Defenders**

B
7.2/10
From $12.50 / $75
Directly competes with Claude Mythos Pre...Lowered refusal boundary on defensive-se...
Updated 2026-09-14
Microsoft MAI-Thinking-1 logo

Microsoft MAI-Thinking-1

Microsoft's first in-house reasoning model -- launched 2026-06-02 at Build as the flagship of seven new MAI models. 35B-active / ~1T-total sparse Mixture-of-Experts, 256K context. AIME 2025 97.0%, matches leading models on SWE-Bench Pro, and beat Claude Sonnet 4.6 in human-preference testing. Available on Microsoft Foundry + OpenRouter / Fireworks / Baseten

B
7.5/10
From Not disclosed
Microsoft's first in-house frontier-clas...Strong published reasoning numbers: AIME...
Updated 2026-06-02
MiMo (Xiaomi) logo

MiMo (Xiaomi)

Xiaomi's MiMo-V2.5 family launched 2026-04-22 -- Pro (1T total / 42B active MoE, 1M context, native vision+audio reasoning), Multimodal base, TTS (3 sub-models: base, VoiceDesign, VoiceClone), and ASR (open-source, English + Chinese + major dialects). Full voice pipeline for the agent era. Extra-charge 1M-context tier removed at launch

A
8.3/10
Free tierFrom $0
Full voice pipeline shipped together: a ...Native multimodal in MiMo-V2.5-Pro is th...
Updated 2026-07-04