Frontier-class LLMs from DeepSeek, Qwen, GLM, Kimi, MiniMax & Doubao — dramatically cheaper than GPT & Claude, mostly open-weight, and ready for your stack. What they cost, how they benchmark, and how to access them from anywhere.
60–90% cheaper; gap narrows further after GPT-5.6 Luna permanent 80% cut ($0.2/$1.2)
at comparable quality
US enterprise token share
Chinese models on OpenRouter: 4.5% → 46% of US-company tokens (2025→2026)
on OpenRouter, 2025 → 2026
Performance gap
Roughly 6–9 months behind · near parity on most tasks
and closing every quarter
Open weights
多数中国前沿模型开放权重(MIT / Apache 2.0),可本地自部署;开源与闭源 API 两条线并行
self-host most frontier models
15 shown
Model Index
SPECIMEN INDEX15 ENTRIES
🇨🇳FlagshipOpen · MIT
DeepSeek-V4-Pro
DeepSeek
DeepSeek's flagship LLM released in April 2026. MoE architecture with 1.6T total parameters and 49B active, native 1M-token context, fully open-sourced under the MIT License. Scores 80.6% on SWE-bench Verified and 3206 on Codeforces, matching the world's top closed-source models.
Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.
🌍FrontierClosed
Claude Fable 5
Anthropic
Anthropic's Mythos-class flagship model released on June 9, 2026 — the first Mythos-class model available to the public. Leads substantially in coding, long-horizon agent tasks and complex reasoning. Supports Adaptive Thinking and can run complex task chains continuously for hours.
Context
1M Tokens
Max Output
128K Tokens
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$50USD
Input $10
※ Official API pricing (anthropic.com, Sep 2026): input $10 / cache read $0.25 / output $50 per M tokens, cache write $12.5; ~2x Opus; in Max/Team Premium since Jul 20, 2026; Batch 50% off
Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning
How to access
Direct API; included in Max/Team Premium subscriptions.
🌍FlagshipClosed
GPT-5.5
OpenAI
OpenAI's 2026 flagship model. 1.05M-token context, 130K max output, native multimodal (text/image/audio), leading multiple benchmarks.
Context
1.05M Tokens
Max Output
130,000 Tokens
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$30USD
Input $5
Professional codingMultimodal understanding & generationEnterprise-grade reasoning
How to access
Direct API at platform.openai.com. Not directly available in mainland China.
🇨🇳BudgetClosed
Doubao-Seed-2.0-Pro(豆包2.0 Pro)
ByteDance / Volcano Engine
ByteDance's flagship LLM released on February 14, 2026. Gold-medal level in the IMO/CMO math olympiads and ICPC programming contest, surpassing Gemini 3 Pro on the Putnam benchmark. Omni-modal, strong at scientific reasoning and agents, with over 100 million monthly active users.
13×cheaper than GPT-5.5 (output $/M)
Context
256K Tokens
Max Output
—
Parameters
Not disclosed
Availability
China (Volcano Engine)
Output / 1M tokens$2.35USD
Input $0.47
※ Lite: ¥0.6 in / ¥3.6 out; Mini: ¥0.2 in / ¥2 out
Via ByteDance Volcano Engine (火山引擎). Primarily China; limited global access.
🇨🇳FlagshipOpen · Open (weights releasing Jul 27, 2026)
Kimi K3
Moonshot AI
Moonshot AI's 2.8T-parameter frontier flagship released in July 2026. 896-expert MoE with 16 experts active per token, and a flat-rate 1M context with no tiered surcharge. KDA hybrid attention delivers a 6.3x decode speedup; frontend coding ranks #1 globally, matching Claude Opus on coding & reasoning at half the cost. Open weights released July 27, 2026.
2×cheaper than GPT-5.5 (output $/M)
Context
1M Tokens
Max Output
128K Tokens
Parameters
2.8T MoE (16 experts active)
Availability
Global
Output / 1M tokens$14.71USD
Input $2.94
※ Official pay-as-you-go pricing (platform.kimi.com, Aug 2026): input cache-miss ¥20 / cache-hit ¥2 / output ¥100 per M tokens, single 1M-context tier; automatic context caching enabled (prefix cache hits when prompt >256 tokens); no thinking/non-thinking price split
Frontend_coding #1 globallyArtificial_Analysis_Index Top tier
API at platform.moonshot.cn and kimi.ai. OpenAI SDK-compatible — zero migration cost. Open weights planned Jul 27, 2026.
🇨🇳FlagshipOpen · MIT
GLM-5.2
Zhipu AI
Z.ai (Zhipu)'s 2026 open flagship, MIT-licensed open weights on Hugging Face. Matches Claude Opus 4.8 at ~80% lower cost — overall spend ≈ 20% of Claude Opus. First-week API calls jumped 27x, praised by prominent engineers as matching US systems at a fraction of the cost.
7×cheaper than GPT-5.5 (output $/M)
Context
200K Tokens
Max Output
128K Tokens
Parameters
MoE
Availability
Global
Output / 1M tokens$4.12USD
Input $1.18
※ Official pay-as-you-go pricing (bigmodel.cn, Aug 2026): input ¥8 / output ¥28 / cache hit ¥2 per M tokens; cache storage free for a limited time
vs_Claude_Opus_4.8 ~parity, 80% cheaper
General reasoningCodingCost-cutting Claude replacement
How to access
Open weights on Hugging Face. API at bigmodel.cn and on OpenRouter.
🇨🇳BudgetOpen · Apache 2.0
MiniMax M2.5
MiniMax
MiniMax's efficient workhorse, consistently in the global top three by API call volume. Excellent price-to-performance for consumer-scale apps.
24×cheaper than GPT-5.5 (output $/M)
Context
200K Tokens
Max Output
64K Tokens
Parameters
MoE
Availability
Global
Output / 1M tokens$1.24USD
Input $0.31
※ Official pricing (platform.minimaxi.com, Aug 2026, listed under legacy models): input ¥2.1 / cache read ¥0.21 / output ¥8.4 per M tokens, single tier, no 50% off (that is M3-only); highspeed edition ¥4.2/¥16.8; top global by call volume
Global call volume Top 3
Roleplay & chatHigh-throughput agentsMultimodal
How to access
API at platform.minimaxi.com (international). Also on OpenRouter.
🌍FlagshipClosed
Claude Opus 5
Anthropic
Released July 24, 2026. Near-Fable-5 performance at half the cost. 2x+ Frontier-Bench vs Opus 4.8. Default model on Claude Max. Supports effort ladder and Fast mode (2.5x speed).
Context
1M Tokens
Max Output
128K
Parameters
Not disclosed
Availability
Global (excluding mainland China)
Output / 1M tokens$25USD
Input $5
※ Same price as Opus 4.8; Fast mode $10/$50; Batch 50% off; cache hits $0.50/MTok (90% savings), cache write $6.25/MTok
anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.
🇨🇳FlagshipOpen
GLM-5.3
Zhipu AI
Z.ai's Aug 14, 2026 flagship: 743B params on the same base as GLM-5.2, uplifted purely by post-training scaling (IndexShare, SAO and the open-sourced Slime RL framework). Coding +50% over GLM-5.2; open-source #1 on Terminal Bench 3.0 and Agents' Last Exam (CLI); agent performance on par with Kimi K3, Claude Opus 4.8 and Claude Fable 5. Live in ZCode, AutoClaw and GLM Coding Plan; full weights in two weeks, API coming soon.
7×cheaper than GPT-5.5 (output $/M)
Context
200K Tokens
Max Output
128K Tokens
Parameters
743B MoE
Availability
全球
Output / 1M tokens$4.12USD
Input $1.18
※ Official pay-as-you-go pricing (bigmodel.cn, Aug 2026): input ¥8 / output ¥28 / cache hit ¥2 per M tokens, same as GLM-5.2; API coming soon, currently usable via GLM Coding Plan subscriptions
AI coding (ZCode / Claude Code platforms)Long-horizon agent tasksComplex software engineering
How to access
Live in ZCode / AutoClaw / GLM Coding Plan; weights open in 2 weeks, API coming soon.
🌍FlagshipClosed
GPT-5.6 Luna
OpenAI
OpenAI's new frontier model family released July 9, 2026, in three tiers: flagship Sol, workhorse Terra and lightweight Luna. Luna targets low-cost high-throughput with ~62% fewer factual errors than GPT-5.5 Instant. Permanent price cuts from July 30: Luna -80%, Terra -20%; Sol gains a Fast mode (2.5x speed at 2x price).
Context
—
Max Output
—
Parameters
Not disclosed
Availability
Global (not mainland China)
Output / 1M tokens$1.2USD
Input $0.2
※ Luna pricing; permanent cuts from 2026-07-30: Terra $2/$12 (-20%), Luna $0.2/$1.2 (-80%); Sol Fast mode 2x price for 2.5x speed
Artificial Analysis Luna beats DeepSeek V4-Pro on cost-efficiency
Batch text processingHigh-throughput agentsEnterprise agents
How to access
Direct API at platform.openai.com. Not directly available in mainland China.
🇨🇳OpenOpen
Tencent Hunyuan Hy4 preview
Tencent
Tencent's new open-source flagship LLM released on August 28, 2026. MoE architecture with 770B total / 49B active parameters and a 1M-token context (960K max input, 64K max output), built for 'planning, executing and reliably delivering' agent capabilities. An open-source Preview build (not GA), live on Tencent Cloud TokenHub and OpenRouter.
11×cheaper than GPT-5.5 (output $/M)
Context
1M Tokens
Max Output
64K Tokens
Parameters
770B MoE (49B active)
Availability
Global (open weights)
Output / 1M tokens$2.65USD
Input $0.88
※ Official launch list price (China site, 2026-08-28): input ¥6 / cache hit ¥0.3 / output ¥18 per M tokens; Tencent Cloud International (Singapore) same model $0.834/$2.501/cache $0.042; early Preview build, not GA
Open weights on Hugging Face / GitHub (Tencent-Hunyuan/Hy4-preview); API via Tencent Cloud TokenHub (console.cloud.tencent.com/tokenhub) pay-as-you-go, China ¥6/¥18; International (Singapore) ~$0.834/$2.501; also on OpenRouter.
🌍FrontierClosed
Claude Fable 5.1
Anthropic
Anthropic's point release on top of Fable 5, released in early September 2026. Retains the Mythos-class flagship capabilities — leading in coding, long-horizon agent tasks and complex reasoning, with Adaptive Thinking for hours-long task chains; 5.1 focuses on long-context stability, tool-call reliability and instruction following.
Context
1M Tokens
Max Output
128K Tokens
Parameters
Not disclosed
Availability
Global (excl. mainland China)
Output / 1M tokens$50USD
Input $10
※ Reuses Claude Fable 5 official API pricing (anthropic.com, Sep 2026): input $10 / cache read $0.25 / output $50 per M tokens, cache write $12.5; Batch 50% off
Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning
How to access
Direct API; included in Max/Team Premium subscription.
🌍FrontierClosed
GPT-6 Astra
OpenAI
OpenAI's next-generation flagship model (codename Astra), released September 2026. First to reach the 'critical' cybersecurity capability tier in internal evals (ExploitBench perfect on known CVEs, autonomously found real zero-days). Native multimodal (text/image/audio); can solve hard math problems like high-dimensional sphere packing and group theory with Lean formal verification. Advanced safety-gated capabilities are review-user only.
Context
2M Tokens
Max Output
200,000 Tokens
Parameters
Not disclosed
Availability
Global (excl. mainland China)
Output / 1M tokens$1.76USD
Input $0.44
※ Official pay-as-you-go pricing (platform.qianwenai.com, Aug 2026): input ¥3 / output ¥12 per M tokens for input <=256K; 256K-1M tier ¥9/¥36; 20% off current promo, context caching supported
Frontier math & formal proofsHigh-assurance agentic tasksComplex multi-step reasoning
How to access
ChatGPT Plus/Pro, API (platform.openai.com); advanced safety capabilities review-user only.
🌍FlagshipClosed
Claude Opus 5.1
Anthropic
Anthropic's September 2026 update to Opus 5, with 2x+ Frontier-Bench performance vs Opus 5. Supports new 'Adaptive Reasoning' mode for complex multi-step tasks. Default model on Claude Max.
Context
1M Tokens
Max Output
128K
Parameters
Not disclosed
Availability
Global (excluding mainland China)
Output / 1M tokens$25USD
Input $5
※ Same price as Opus 5; Adaptive Reasoning mode available; Batch 50% off; cache hits $0.50/MTok (90% savings), cache write $6.25/MTok
anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.
🇨🇳FlagshipOpen · MIT
DeepSeek-V4.1-Flash
DeepSeek
DeepSeek's new-generation Flash model, officially released on September 10, 2026. It uses an entirely new model structure with native multimodal vision understanding and a 1M-token context, and drastically shrinks the KV cache (HBM demand down to 1/4 and SSD down to 1/8 versus the previous generation), cutting the cost of cached Agent workloads. Community tests measure 300–500 tokens/s. DeepSeek says it beats V4 Pro across performance, speed, cost and total latency, and it has taken over V4 Pro API traffic. The weights are open-sourced under the MIT License, enabling local deployment and fine-tuning.
※ New Flash pricing effective 2026-09-10 12:00 (Beijing): off-peak input ¥1 (cache miss) / ¥0.02 (cache hit) / output ¥4 per million tokens; peak = 2× on weekdays 9-12 & 14-18 Beijing time; API model name deepseek-flash, thinking mode on by default
Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.
Why global developers are switching to Chinese AI models
Since 2025, Chinese LLMs have surged from 4.5% to over 46% of US enterprise tokens on OpenRouter. The driver is simple economics: models like DeepSeek V4 and GLM-5.2 deliver near-frontier quality at 60–90% lower cost than GPT-5.5 or Claude Opus. Companies like Lindy, Coinbase, DoorDash and Siemens have already made the switch.
How to choose the right Chinese LLM
For coding agents and long-context reasoning, DeepSeek V4 Pro leads. For self-hosting and multilingual apps, Qwen3 Max offers the strongest open-weight ecosystem. GLM-5.2 is the cheapest Claude replacement, Kimi K2.5 owns the 2M-token long-context niche, and MiniMax M2.5 / Doubao 2.0 win on raw throughput cost. Filter by origin, open weights, or price above.
How to access Chinese AI models from outside China
Most Chinese frontier models are available globally without a Chinese payment method. DeepSeek, Qwen, GLM and MiniMax all accept international cards or run on OpenRouter and Together.ai. Qwen, DeepSeek and GLM also ship open weights on Hugging Face for full self-hosting. Each card above lists exact access steps.