CHINA AI FIELD GUIDE · FOR GLOBAL BUILDERS

CHINESE AI.

MODELS, DECODED

Frontier-class LLMs from DeepSeek, Qwen, GLM, Kimi, MiniMax & Doubao — dramatically cheaper than GPT & Claude, mostly open-weight, and ready for your stack. What they cost, how they benchmark, and how to access them from anywhere.

Prices in USD / 1M tokens
WHY IT MATTERSLIVE SHIFT

Cheaper than US models

60–90% cheaper; gap narrows further after GPT-5.6 Luna permanent 80% cut ($0.2/$1.2)

at comparable quality

US enterprise token share

Chinese models on OpenRouter: 4.5% → 46% of US-company tokens (2025→2026)

on OpenRouter, 2025 → 2026

Performance gap

Roughly 6–9 months behind · near parity on most tasks

and closing every quarter

Open weights

多数中国前沿模型开放权重(MIT / Apache 2.0),可本地自部署;开源与闭源 API 两条线并行

self-host most frontier models

Model Index

SPECIMEN INDEX
15 ENTRIES
🇨🇳FlagshipOpen · MIT

DeepSeek-V4-Pro

DeepSeek

DeepSeek's flagship LLM released in April 2026. MoE architecture with 1.6T total parameters and 49B active, native 1M-token context, fully open-sourced under the MIT License. Scores 80.6% on SWE-bench Verified and 3206 on Codeforces, matching the world's top closed-source models.

cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

384K Tokens

Parameters

1.6T MoE (49B active)

Availability

Global

Output / 1M tokensOff-peak$3.97USD
Input $0.66
⚡ Peak Input $1.32 · Output $3.979:00-12:00 / 14:00-18:00

※ Peak/off-peak billing since Aug 2026: off-peak at half price (peak hours 9-12 & 14-18 Beijing time); cache-hit input ¥0.15 off-peak / ¥0.30 peak

SWE-bench_Verified 80.6%Codeforces 3206
Complex logical reasoningLarge-scale code engineeringMillion-level long-document understanding

How to access

Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.

MITCN
🌍FrontierClosed

Claude Fable 5

Anthropic

Anthropic's Mythos-class flagship model released on June 9, 2026 — the first Mythos-class model available to the public. Leads substantially in coding, long-horizon agent tasks and complex reasoning. Supports Adaptive Thinking and can run complex task chains continuously for hours.

Context

1M Tokens

Max Output

128K Tokens

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$50USD
Input $10

※ Official API pricing (anthropic.com, Sep 2026): input $10 / cache read $0.25 / output $50 per M tokens, cache write $12.5; ~2x Opus; in Max/Team Premium since Jul 20, 2026; Batch 50% off

Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning

How to access

Direct API; included in Max/Team Premium subscriptions.

ProprietaryWEST
🌍FlagshipClosed

GPT-5.5

OpenAI

OpenAI's 2026 flagship model. 1.05M-token context, 130K max output, native multimodal (text/image/audio), leading multiple benchmarks.

Context

1.05M Tokens

Max Output

130,000 Tokens

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$30USD
Input $5
Professional codingMultimodal understanding & generationEnterprise-grade reasoning

How to access

Direct API at platform.openai.com. Not directly available in mainland China.

ProprietaryWEST
🇨🇳BudgetClosed

Doubao-Seed-2.0-Pro(豆包2.0 Pro)

ByteDance / Volcano Engine

ByteDance's flagship LLM released on February 14, 2026. Gold-medal level in the IMO/CMO math olympiads and ICPC programming contest, surpassing Gemini 3 Pro on the Putnam benchmark. Omni-modal, strong at scientific reasoning and agents, with over 100 million monthly active users.

13×cheaper than GPT-5.5 (output $/M)

Context

256K Tokens

Max Output

Parameters

Not disclosed

Availability

China (Volcano Engine)

Output / 1M tokens$2.35USD
Input $0.47

※ Lite: ¥0.6 in / ¥3.6 out; Mini: ¥0.2 in / ¥2 out

IMO_CMO Gold medalICPC Gold medal
Math-olympiad-level reasoningMultimodal content creationEveryday conversation (100M+ MAU)

How to access

Via ByteDance Volcano Engine (火山引擎). Primarily China; limited global access.

ProprietaryCN
🇨🇳FlagshipOpen · Open (weights releasing Jul 27, 2026)

Kimi K3

Moonshot AI

Moonshot AI's 2.8T-parameter frontier flagship released in July 2026. 896-expert MoE with 16 experts active per token, and a flat-rate 1M context with no tiered surcharge. KDA hybrid attention delivers a 6.3x decode speedup; frontend coding ranks #1 globally, matching Claude Opus on coding & reasoning at half the cost. Open weights released July 27, 2026.

cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

128K Tokens

Parameters

2.8T MoE (16 experts active)

Availability

Global

Output / 1M tokens$14.71USD
Input $2.94

※ Official pay-as-you-go pricing (platform.kimi.com, Aug 2026): input cache-miss ¥20 / cache-hit ¥2 / output ¥100 per M tokens, single 1M-context tier; automatic context caching enabled (prefix cache hits when prompt >256 tokens); no thinking/non-thinking price split

Frontend_coding #1 globallyArtificial_Analysis_Index Top tier
Coding agentsLong-context reasoningComplex multi-step tasks

How to access

API at platform.moonshot.cn and kimi.ai. OpenAI SDK-compatible — zero migration cost. Open weights planned Jul 27, 2026.

Open (weights releasing Jul 27, 2026)CN
🇨🇳FlagshipOpen · MIT

GLM-5.2

Zhipu AI

Z.ai (Zhipu)'s 2026 open flagship, MIT-licensed open weights on Hugging Face. Matches Claude Opus 4.8 at ~80% lower cost — overall spend ≈ 20% of Claude Opus. First-week API calls jumped 27x, praised by prominent engineers as matching US systems at a fraction of the cost.

cheaper than GPT-5.5 (output $/M)

Context

200K Tokens

Max Output

128K Tokens

Parameters

MoE

Availability

Global

Output / 1M tokens$4.12USD
Input $1.18

※ Official pay-as-you-go pricing (bigmodel.cn, Aug 2026): input ¥8 / output ¥28 / cache hit ¥2 per M tokens; cache storage free for a limited time

vs_Claude_Opus_4.8 ~parity, 80% cheaper
General reasoningCodingCost-cutting Claude replacement

How to access

Open weights on Hugging Face. API at bigmodel.cn and on OpenRouter.

MITCN
🇨🇳BudgetOpen · Apache 2.0

MiniMax M2.5

MiniMax

MiniMax's efficient workhorse, consistently in the global top three by API call volume. Excellent price-to-performance for consumer-scale apps.

24×cheaper than GPT-5.5 (output $/M)

Context

200K Tokens

Max Output

64K Tokens

Parameters

MoE

Availability

Global

Output / 1M tokens$1.24USD
Input $0.31

※ Official pricing (platform.minimaxi.com, Aug 2026, listed under legacy models): input ¥2.1 / cache read ¥0.21 / output ¥8.4 per M tokens, single tier, no 50% off (that is M3-only); highspeed edition ¥4.2/¥16.8; top global by call volume

Global call volume Top 3
Roleplay & chatHigh-throughput agentsMultimodal

How to access

API at platform.minimaxi.com (international). Also on OpenRouter.

Apache 2.0CN
🌍FlagshipClosed

Claude Opus 5

Anthropic

Released July 24, 2026. Near-Fable-5 performance at half the cost. 2x+ Frontier-Bench vs Opus 4.8. Default model on Claude Max. Supports effort ladder and Fast mode (2.5x speed).

Context

1M Tokens

Max Output

128K

Parameters

Not disclosed

Availability

Global (excluding mainland China)

Output / 1M tokens$25USD
Input $5

※ Same price as Opus 4.8; Fast mode $10/$50; Batch 50% off; cache hits $0.50/MTok (90% savings), cache write $6.25/MTok

DeepSWE v1.1 68.8%FrontierCode v1.1 53.4%
Complex agentic codingEnterprise code reviewLong-horizon autonomous tasks

How to access

anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.

ProprietaryWEST
🇨🇳FlagshipOpen

GLM-5.3

Zhipu AI

Z.ai's Aug 14, 2026 flagship: 743B params on the same base as GLM-5.2, uplifted purely by post-training scaling (IndexShare, SAO and the open-sourced Slime RL framework). Coding +50% over GLM-5.2; open-source #1 on Terminal Bench 3.0 and Agents' Last Exam (CLI); agent performance on par with Kimi K3, Claude Opus 4.8 and Claude Fable 5. Live in ZCode, AutoClaw and GLM Coding Plan; full weights in two weeks, API coming soon.

cheaper than GPT-5.5 (output $/M)

Context

200K Tokens

Max Output

128K Tokens

Parameters

743B MoE

Availability

全球

Output / 1M tokens$4.12USD
Input $1.18

※ Official pay-as-you-go pricing (bigmodel.cn, Aug 2026): input ¥8 / output ¥28 / cache hit ¥2 per M tokens, same as GLM-5.2; API coming soon, currently usable via GLM Coding Plan subscriptions

Terminal_Bench_3.0 Open-source #1Agents_Last_Exam_CLI Open-source #1
AI coding (ZCode / Claude Code platforms)Long-horizon agent tasksComplex software engineering

How to access

Live in ZCode / AutoClaw / GLM Coding Plan; weights open in 2 weeks, API coming soon.

CN
🌍FlagshipClosed

GPT-5.6 Luna

OpenAI

OpenAI's new frontier model family released July 9, 2026, in three tiers: flagship Sol, workhorse Terra and lightweight Luna. Luna targets low-cost high-throughput with ~62% fewer factual errors than GPT-5.5 Instant. Permanent price cuts from July 30: Luna -80%, Terra -20%; Sol gains a Fast mode (2.5x speed at 2x price).

Context

Max Output

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$1.2USD
Input $0.2

※ Luna pricing; permanent cuts from 2026-07-30: Terra $2/$12 (-20%), Luna $0.2/$1.2 (-80%); Sol Fast mode 2x price for 2.5x speed

Artificial Analysis Luna beats DeepSeek V4-Pro on cost-efficiency
Batch text processingHigh-throughput agentsEnterprise agents

How to access

Direct API at platform.openai.com. Not directly available in mainland China.

WEST
🇨🇳OpenOpen

Tencent Hunyuan Hy4 preview

Tencent

Tencent's new open-source flagship LLM released on August 28, 2026. MoE architecture with 770B total / 49B active parameters and a 1M-token context (960K max input, 64K max output), built for 'planning, executing and reliably delivering' agent capabilities. An open-source Preview build (not GA), live on Tencent Cloud TokenHub and OpenRouter.

11×cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

64K Tokens

Parameters

770B MoE (49B active)

Availability

Global (open weights)

Output / 1M tokens$2.65USD
Input $0.88

※ Official launch list price (China site, 2026-08-28): input ¥6 / cache hit ¥0.3 / output ¥18 per M tokens; Tencent Cloud International (Singapore) same model $0.834/$2.501/cache $0.042; early Preview build, not GA

Complex agent tasksLong-document/code understandingEnterprise reasoning

How to access

Open weights on Hugging Face / GitHub (Tencent-Hunyuan/Hy4-preview); API via Tencent Cloud TokenHub (console.cloud.tencent.com/tokenhub) pay-as-you-go, China ¥6/¥18; International (Singapore) ~$0.834/$2.501; also on OpenRouter.

CN
🌍FrontierClosed

Claude Fable 5.1

Anthropic

Anthropic's point release on top of Fable 5, released in early September 2026. Retains the Mythos-class flagship capabilities — leading in coding, long-horizon agent tasks and complex reasoning, with Adaptive Thinking for hours-long task chains; 5.1 focuses on long-context stability, tool-call reliability and instruction following.

Context

1M Tokens

Max Output

128K Tokens

Parameters

Not disclosed

Availability

Global (excl. mainland China)

Output / 1M tokens$50USD
Input $10

※ Reuses Claude Fable 5 official API pricing (anthropic.com, Sep 2026): input $10 / cache read $0.25 / output $50 per M tokens, cache write $12.5; Batch 50% off

Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning

How to access

Direct API; included in Max/Team Premium subscription.

ProprietaryWEST
🌍FrontierClosed

GPT-6 Astra

OpenAI

OpenAI's next-generation flagship model (codename Astra), released September 2026. First to reach the 'critical' cybersecurity capability tier in internal evals (ExploitBench perfect on known CVEs, autonomously found real zero-days). Native multimodal (text/image/audio); can solve hard math problems like high-dimensional sphere packing and group theory with Lean formal verification. Advanced safety-gated capabilities are review-user only.

Context

2M Tokens

Max Output

200,000 Tokens

Parameters

Not disclosed

Availability

Global (excl. mainland China)

Output / 1M tokens$1.76USD
Input $0.44

※ Official pay-as-you-go pricing (platform.qianwenai.com, Aug 2026): input ¥3 / output ¥12 per M tokens for input <=256K; 256K-1M tier ¥9/¥36; 20% off current promo, context caching supported

Frontier math & formal proofsHigh-assurance agentic tasksComplex multi-step reasoning

How to access

ChatGPT Plus/Pro, API (platform.openai.com); advanced safety capabilities review-user only.

ProprietaryWEST
🌍FlagshipClosed

Claude Opus 5.1

Anthropic

Anthropic's September 2026 update to Opus 5, with 2x+ Frontier-Bench performance vs Opus 5. Supports new 'Adaptive Reasoning' mode for complex multi-step tasks. Default model on Claude Max.

Context

1M Tokens

Max Output

128K

Parameters

Not disclosed

Availability

Global (excluding mainland China)

Output / 1M tokens$25USD
Input $5

※ Same price as Opus 5; Adaptive Reasoning mode available; Batch 50% off; cache hits $0.50/MTok (90% savings), cache write $6.25/MTok

FrontierCode v1.1 55.8%
Complex agentic codingEnterprise code reviewLong-horizon autonomous tasks

How to access

anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.

ProprietaryWEST
🇨🇳FlagshipOpen · MIT

DeepSeek-V4.1-Flash

DeepSeek

DeepSeek's new-generation Flash model, officially released on September 10, 2026. It uses an entirely new model structure with native multimodal vision understanding and a 1M-token context, and drastically shrinks the KV cache (HBM demand down to 1/4 and SSD down to 1/8 versus the previous generation), cutting the cost of cached Agent workloads. Community tests measure 300–500 tokens/s. DeepSeek says it beats V4 Pro across performance, speed, cost and total latency, and it has taken over V4 Pro API traffic. The weights are open-sourced under the MIT License, enabling local deployment and fine-tuning.

51×cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

384K Tokens

Parameters

Not disclosed

Availability

Global

Output / 1M tokensOff-peak$0.59USD
Input $0.15
⚡ Peak Input $0.29 · Output $1.18Mon-Fri 9:00-12:00 / 14:00-18:00

※ New Flash pricing effective 2026-09-10 12:00 (Beijing): off-peak input ¥1 (cache miss) / ¥0.02 (cache hit) / output ¥4 per million tokens; peak = 2× on weekdays 9-12 & 14-18 Beijing time; API model name deepseek-flash, thinking mode on by default

High-concurrency everyday chatAgentic task executionVision understanding (screenshots, documents)

How to access

Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.

MITCN

Why global developers are switching to Chinese AI models

Since 2025, Chinese LLMs have surged from 4.5% to over 46% of US enterprise tokens on OpenRouter. The driver is simple economics: models like DeepSeek V4 and GLM-5.2 deliver near-frontier quality at 60–90% lower cost than GPT-5.5 or Claude Opus. Companies like Lindy, Coinbase, DoorDash and Siemens have already made the switch.

How to choose the right Chinese LLM

For coding agents and long-context reasoning, DeepSeek V4 Pro leads. For self-hosting and multilingual apps, Qwen3 Max offers the strongest open-weight ecosystem. GLM-5.2 is the cheapest Claude replacement, Kimi K2.5 owns the 2M-token long-context niche, and MiniMax M2.5 / Doubao 2.0 win on raw throughput cost. Filter by origin, open weights, or price above.

How to access Chinese AI models from outside China

Most Chinese frontier models are available globally without a Chinese payment method. DeepSeek, Qwen, GLM and MiniMax all accept international cards or run on OpenRouter and Together.ai. Qwen, DeepSeek and GLM also ship open weights on Hugging Face for full self-hosting. Each card above lists exact access steps.