Frontier-class LLMs from DeepSeek, Qwen, GLM, Kimi, MiniMax & Doubao — dramatically cheaper than GPT & Claude, mostly open-weight, and ready for your stack. What they cost, how they benchmark, and how to access them from anywhere.
Chinese models on OpenRouter: 4.5% → 46% of US enterprise tokens (2025→2026)
on OpenRouter, 2025 → 2026
Performance gap
~6–9 months behind · near parity on most tasks
and closing every quarter
Open weights
Most Chinese frontier models ship open weights (MIT / Apache 2.0)
self-host most frontier models
14 shown
Model Index
SPECIMEN INDEX14 ENTRIES
🇨🇳FlagshipOpen · MIT
DeepSeek-V4-Pro
DeepSeek
DeepSeek's flagship LLM released in April 2026. MoE architecture with 1.6T total parameters and 49B active, native 1M-token context, fully open-sourced under the MIT License. Scores 80.6% on SWE-bench Verified and 3206 on Codeforces, matching the world's top closed-source models.
34×cheaper than GPT-5.5 (output $/M)
Context
1M Tokens
Max Output
384K Tokens
Parameters
1.6T MoE (49B active)
Availability
Global
Output / 1M tokens$0.88USD
Input $0.44
※ Permanent 75% discount since May 2026; peak/off-peak pricing from July: weekday 9-12/14-18 BJT doubles (peak ¥6/¥12)
Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.
🇨🇳BudgetOpen · MIT
DeepSeek-V4-Flash
DeepSeek
The economy edition of the DeepSeek V4 series. 284B total parameters, 13B active, focused on high concurrency and extreme cost efficiency, 1M context, open-sourced under the MIT License.
102×cheaper than GPT-5.5 (output $/M)
Context
1M Tokens
Max Output
—
Parameters
284B MoE (13B active)
Availability
Global
Output / 1M tokens$0.29USD
Input $0.15
※ Peak/off-peak from July: weekday 9-12/14-18 BJT doubles (peak ¥2/¥4)
High-concurrency everyday chatLightweight text generationCost-sensitive batch processing
How to access
Same as DeepSeek V4 Pro — direct API, OpenRouter, Together.ai.
🌍FrontierClosed
Claude Fable 5
Anthropic
Anthropic's Mythos-class flagship model released on June 9, 2026 — the first Mythos-class model available to the public. Leads substantially in coding, long-horizon agent tasks and complex reasoning. Supports Adaptive Thinking and can run complex task chains continuously for hours.
Context
1M Tokens
Max Output
128K Tokens
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$50USD
Input $10
※ About twice Opus; included in Max/Team Premium subscriptions since July 20, 2026
Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning
How to access
Direct API; included in Max/Team Premium subscriptions.
🌍FlagshipClosed
Claude Opus 4.6
Anthropic
The top reasoning model in Anthropic's Claude 4 series. 80.9% on SWE-bench, 1M-token context — the first choice for enterprise-grade deep reasoning and code review.
Direct API at anthropic.com. Not directly available in mainland China.
🌍FlagshipClosed
GPT-5.5
OpenAI
OpenAI's 2026 flagship model. 1.05M-token context, 130K max output, native multimodal (text/image/audio), leading multiple benchmarks.
Context
1.05M Tokens
Max Output
130,000 Tokens
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$30USD
Input $5
Professional codingMultimodal understanding & generationEnterprise-grade reasoning
How to access
Direct API at platform.openai.com. Not directly available in mainland China.
🌍FlagshipClosed
Gemini 3.0 Pro
Google Dee
Google Gemini 3's flagship reasoning model. Scores 77.1% on ARC-AGI-2 and supports a deep-thinking mode.
Context
1M Tokens
Max Output
—
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$12USD
Input $2
ARC-AGI-2 77.1%
Complex scientific reasoningMulti-step problem solvingLong-video understanding
How to access
Direct API at ai.google.dev. Not directly available in mainland China.
🇨🇳BudgetClosed
Doubao-Seed-2.0-Pro(豆包2.0 Pro)
ByteDance / Volcano Engine
ByteDance's flagship LLM released on February 14, 2026. Gold-medal level in the IMO/CMO math olympiads and ICPC programming contest, surpassing Gemini 3 Pro on the Putnam benchmark. Omni-modal, strong at scientific reasoning and agents, with over 100 million monthly active users.
13×cheaper than GPT-5.5 (output $/M)
Context
256K Tokens
Max Output
—
Parameters
Not disclosed
Availability
China (Volcano Engine)
Output / 1M tokens$2.35USD
Input $0.47
※ Lite: ¥0.6 in / ¥3.6 out; Mini: ¥0.2 in / ¥2 out
Via ByteDance Volcano Engine (火山引擎). Primarily China; limited global access.
🌍OpenOpen · Meta Llama License
Llama 4 Scout / Maverick
Meta
Meta's open-source MoE multimodal model series released in April 2025. Scout: 109B total / 17B active with a 10M ultra-long context (the industry's longest); Maverick: 400B total / 17B active with 1M context. Native multimodal, supporting 12+ languages.
Context
10M Tokens
Max Output
—
Parameters
Scout 109B / Maverick 400B MoE (17B active)
Availability
Global (open weights)
Output / 1M tokens$0.6USD
Input $0.2
※ Open weights, free to self-host; cloud hosts charge by compute (AWS Bedrock Scout ~$0.17/$0.66, Maverick ~$0.24/$0.97); hosted reference price shown
Open weights on Hugging Face. Hosted via Together.ai, Fireworks, Groq.
🇨🇳FlagshipOpen · Open (weights releasing Jul 27, 2026)
Kimi K3
Moonshot AI
Moonshot AI's 2.8T-parameter frontier flagship released in July 2026. 896-expert MoE with 16 experts active per token, and a flat-rate 1M context with no tiered surcharge. KDA hybrid attention delivers a 6.3x decode speedup; frontend coding ranks #1 globally, matching Claude Opus on coding & reasoning at half the cost. Open weights released July 27, 2026.
2×cheaper than GPT-5.5 (output $/M)
Context
1M Tokens
Max Output
128K Tokens
Parameters
2.8T MoE (16 experts active)
Availability
Global
Output / 1M tokens$15.88USD
Input $3.18
※ Cache-hit input only $0.30/M (90%+ cache rate in coding); full 1M context at flat rate; the most expensive Chinese model, still ~50% cheaper than the Western frontier
Frontend_coding #1 globallyArtificial_Analysis_Index Top tier
API at platform.moonshot.cn and kimi.ai. OpenAI SDK-compatible — zero migration cost. Open weights planned Jul 27, 2026.
🇨🇳FlagshipOpen · MIT
GLM-5.2
Zhipu AI
Z.ai (Zhipu)'s 2026 open flagship, MIT-licensed open weights on Hugging Face. Matches Claude Opus 4.8 at ~80% lower cost — overall spend ≈ 20% of Claude Opus. First-week API calls jumped 27x, praised by prominent engineers as matching US systems at a fraction of the cost.
35×cheaper than GPT-5.5 (output $/M)
Context
200K Tokens
Max Output
128K Tokens
Parameters
MoE
Availability
Global
Output / 1M tokens$0.85USD
Input $0.21
※ Overall cost ≈ 20% of Claude Opus; first-week calls jumped 27x
vs_Claude_Opus_4.8 ~parity, 80% cheaper
General reasoningCodingCost-cutting Claude replacement
How to access
Open weights on Hugging Face. API at bigmodel.cn and on OpenRouter.
🇨🇳FlagshipOpen · Apache 2.0
Qwen3 Max
Alibaba
Alibaba's open-weight champion and the most-downloaded open model family on Earth. Strongest Chinese-language and agentic performance among open models.
36×cheaper than GPT-5.5 (output $/M)
Context
256K Tokens
Max Output
32K Tokens
Parameters
235B MoE (22B active)
Availability
Global
Output / 1M tokens$0.83USD
Input $0.28
※ Open-weight family with 1B–235B sizes; ~1B downloads (50%+ of global open-model downloads).
MMLU (Chinese) 88.1HumanEval 82.4
Self-hostingMultilingual appsAgents & tool use
How to access
Open weights on Hugging Face & ModelScope. Hosted API via Alibaba DashScope (global) or OpenRouter. Self-host freely.
🇨🇳FlagshipOpen · Modified MIT
Kimi K2.5
Moonshot A
Moonshot's long-context powerhouse with a 2M-token window — the go-to for ingesting entire codebases or document libraries in one call.
API at platform.moonshot.cn (China). Globally accessible via OpenRouter.
🇨🇳BudgetOpen · Apache 2.0
MiniMax M2.5
MiniMax
MiniMax's efficient workhorse, consistently in the global top three by API call volume. Excellent price-to-performance for consumer-scale apps.
75×cheaper than GPT-5.5 (output $/M)
Context
200K Tokens
Max Output
64K Tokens
Parameters
MoE
Availability
Global
Output / 1M tokens$0.4USD
Input $0.1
※ Among the cheapest frontier-class models; top-3 global by call volume.
Global call volume Top 3
Roleplay & chatHigh-throughput agentsMultimodal
How to access
API at platform.minimaxi.com (international). Also on OpenRouter.
🌍FlagshipClosed
Claude Opus 5
Anthropic
Released July 24, 2026. Near-Fable-5 performance at half the cost. 2x+ Frontier-Bench vs Opus 4.8. Default model on Claude Max. Supports effort ladder and Fast mode (2.5x speed).
Context
1M Tokens
Max Output
128K
Parameters
Not disclosed
Availability
Global (excluding mainland China)
Output / 1M tokens$25USD
Input $5
※ Same price as Opus 4.8; Fast mode $10/$50; Batch 50% off; cache hits $0.50/MTok (90% savings)
anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.
Why global developers are switching to Chinese AI models
Since 2025, Chinese LLMs have surged from 4.5% to over 46% of US enterprise tokens on OpenRouter. The driver is simple economics: models like DeepSeek V4 and GLM-5.2 deliver near-frontier quality at 60–90% lower cost than GPT-5.5 or Claude Opus. Companies like Lindy, Coinbase, DoorDash and Siemens have already made the switch.
How to choose the right Chinese LLM
For coding agents and long-context reasoning, DeepSeek V4 Pro leads. For self-hosting and multilingual apps, Qwen3 Max offers the strongest open-weight ecosystem. GLM-5.2 is the cheapest Claude replacement, Kimi K2.5 owns the 2M-token long-context niche, and MiniMax M2.5 / Doubao 2.0 win on raw throughput cost. Filter by origin, open weights, or price above.
How to access Chinese AI models from outside China
Most Chinese frontier models are available globally without a Chinese payment method. DeepSeek, Qwen, GLM and MiniMax all accept international cards or run on OpenRouter and Together.ai. Qwen, DeepSeek and GLM also ship open weights on Hugging Face for full self-hosting. Each card above lists exact access steps.