CHINA AI FIELD GUIDE · FOR GLOBAL BUILDERS

CHINESE AI.

MODELS, DECODED

Frontier-class LLMs from DeepSeek, Qwen, GLM, Kimi, MiniMax & Doubao — dramatically cheaper than GPT & Claude, mostly open-weight, and ready for your stack. What they cost, how they benchmark, and how to access them from anywhere.

Prices in USD / 1M tokens
WHY IT MATTERSLIVE SHIFT

Cheaper than US models

60–90% cheaper

at comparable quality

US enterprise token share

Chinese models on OpenRouter: 4.5% → 46% of US enterprise tokens (2025→2026)

on OpenRouter, 2025 → 2026

Performance gap

~6–9 months behind · near parity on most tasks

and closing every quarter

Open weights

Most Chinese frontier models ship open weights (MIT / Apache 2.0)

self-host most frontier models

Model Index

SPECIMEN INDEX
14 ENTRIES
🇨🇳FlagshipOpen · MIT

DeepSeek-V4-Pro

DeepSeek

DeepSeek's flagship LLM released in April 2026. MoE architecture with 1.6T total parameters and 49B active, native 1M-token context, fully open-sourced under the MIT License. Scores 80.6% on SWE-bench Verified and 3206 on Codeforces, matching the world's top closed-source models.

34×cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

384K Tokens

Parameters

1.6T MoE (49B active)

Availability

Global

Output / 1M tokens$0.88USD
Input $0.44

※ Permanent 75% discount since May 2026; peak/off-peak pricing from July: weekday 9-12/14-18 BJT doubles (peak ¥6/¥12)

SWE-bench_Verified 80.6%Codeforces 3206
Complex logical reasoningLarge-scale code engineeringMillion-level long-document understanding

How to access

Direct API at platform.deepseek.com (international cards accepted). Also on OpenRouter & Together.ai — no Chinese payment method required.

MITCN
🇨🇳BudgetOpen · MIT

DeepSeek-V4-Flash

DeepSeek

The economy edition of the DeepSeek V4 series. 284B total parameters, 13B active, focused on high concurrency and extreme cost efficiency, 1M context, open-sourced under the MIT License.

102×cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

Parameters

284B MoE (13B active)

Availability

Global

Output / 1M tokens$0.29USD
Input $0.15

※ Peak/off-peak from July: weekday 9-12/14-18 BJT doubles (peak ¥2/¥4)

High-concurrency everyday chatLightweight text generationCost-sensitive batch processing

How to access

Same as DeepSeek V4 Pro — direct API, OpenRouter, Together.ai.

MITCN
🌍FrontierClosed

Claude Fable 5

Anthropic

Anthropic's Mythos-class flagship model released on June 9, 2026 — the first Mythos-class model available to the public. Leads substantially in coding, long-horizon agent tasks and complex reasoning. Supports Adaptive Thinking and can run complex task chains continuously for hours.

Context

1M Tokens

Max Output

128K Tokens

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$50USD
Input $10

※ About twice Opus; included in Max/Team Premium subscriptions since July 20, 2026

Large codebase migration & refactoringLong-horizon autonomous agentsComplex multi-step reasoning

How to access

Direct API; included in Max/Team Premium subscriptions.

ProprietaryWEST
🌍FlagshipClosed

Claude Opus 4.6

Anthropic

The top reasoning model in Anthropic's Claude 4 series. 80.9% on SWE-bench, 1M-token context — the first choice for enterprise-grade deep reasoning and code review.

Context

1M Tokens

Max Output

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$25USD
Input $5
SWE-bench_Verified 80.9%
Enterprise code reviewLarge codebase refactoringComplex agent planning

How to access

Direct API at anthropic.com. Not directly available in mainland China.

ProprietaryWEST
🌍FlagshipClosed

GPT-5.5

OpenAI

OpenAI's 2026 flagship model. 1.05M-token context, 130K max output, native multimodal (text/image/audio), leading multiple benchmarks.

Context

1.05M Tokens

Max Output

130,000 Tokens

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$30USD
Input $5
Professional codingMultimodal understanding & generationEnterprise-grade reasoning

How to access

Direct API at platform.openai.com. Not directly available in mainland China.

ProprietaryWEST
🌍FlagshipClosed

Gemini 3.0 Pro

Google Dee

Google Gemini 3's flagship reasoning model. Scores 77.1% on ARC-AGI-2 and supports a deep-thinking mode.

Context

1M Tokens

Max Output

Parameters

Not disclosed

Availability

Global (not mainland China)

Output / 1M tokens$12USD
Input $2
ARC-AGI-2 77.1%
Complex scientific reasoningMulti-step problem solvingLong-video understanding

How to access

Direct API at ai.google.dev. Not directly available in mainland China.

ProprietaryWEST
🇨🇳BudgetClosed

Doubao-Seed-2.0-Pro(豆包2.0 Pro)

ByteDance / Volcano Engine

ByteDance's flagship LLM released on February 14, 2026. Gold-medal level in the IMO/CMO math olympiads and ICPC programming contest, surpassing Gemini 3 Pro on the Putnam benchmark. Omni-modal, strong at scientific reasoning and agents, with over 100 million monthly active users.

13×cheaper than GPT-5.5 (output $/M)

Context

256K Tokens

Max Output

Parameters

Not disclosed

Availability

China (Volcano Engine)

Output / 1M tokens$2.35USD
Input $0.47

※ Lite: ¥0.6 in / ¥3.6 out; Mini: ¥0.2 in / ¥2 out

IMO_CMO Gold medalICPC Gold medal
Math-olympiad-level reasoningMultimodal content creationEveryday conversation (100M+ MAU)

How to access

Via ByteDance Volcano Engine (火山引擎). Primarily China; limited global access.

ProprietaryCN
🌍OpenOpen · Meta Llama License

Llama 4 Scout / Maverick

Meta

Meta's open-source MoE multimodal model series released in April 2025. Scout: 109B total / 17B active with a 10M ultra-long context (the industry's longest); Maverick: 400B total / 17B active with 1M context. Native multimodal, supporting 12+ languages.

Context

10M Tokens

Max Output

Parameters

Scout 109B / Maverick 400B MoE (17B active)

Availability

Global (open weights)

Output / 1M tokens$0.6USD
Input $0.2

※ Open weights, free to self-host; cloud hosts charge by compute (AWS Bedrock Scout ~$0.17/$0.66, Maverick ~$0.24/$0.97); hosted reference price shown

Ultra-long document processing (10M context)Multilingual applicationsOn-premise private deployment

How to access

Open weights on Hugging Face. Hosted via Together.ai, Fireworks, Groq.

Meta Llama LicenseWEST
🇨🇳FlagshipOpen · Open (weights releasing Jul 27, 2026)

Kimi K3

Moonshot AI

Moonshot AI's 2.8T-parameter frontier flagship released in July 2026. 896-expert MoE with 16 experts active per token, and a flat-rate 1M context with no tiered surcharge. KDA hybrid attention delivers a 6.3x decode speedup; frontend coding ranks #1 globally, matching Claude Opus on coding & reasoning at half the cost. Open weights released July 27, 2026.

cheaper than GPT-5.5 (output $/M)

Context

1M Tokens

Max Output

128K Tokens

Parameters

2.8T MoE (16 experts active)

Availability

Global

Output / 1M tokens$15.88USD
Input $3.18

※ Cache-hit input only $0.30/M (90%+ cache rate in coding); full 1M context at flat rate; the most expensive Chinese model, still ~50% cheaper than the Western frontier

Frontend_coding #1 globallyArtificial_Analysis_Index Top tier
Coding agentsLong-context reasoningComplex multi-step tasks

How to access

API at platform.moonshot.cn and kimi.ai. OpenAI SDK-compatible — zero migration cost. Open weights planned Jul 27, 2026.

Open (weights releasing Jul 27, 2026)CN
🇨🇳FlagshipOpen · MIT

GLM-5.2

Zhipu AI

Z.ai (Zhipu)'s 2026 open flagship, MIT-licensed open weights on Hugging Face. Matches Claude Opus 4.8 at ~80% lower cost — overall spend ≈ 20% of Claude Opus. First-week API calls jumped 27x, praised by prominent engineers as matching US systems at a fraction of the cost.

35×cheaper than GPT-5.5 (output $/M)

Context

200K Tokens

Max Output

128K Tokens

Parameters

MoE

Availability

Global

Output / 1M tokens$0.85USD
Input $0.21

※ Overall cost ≈ 20% of Claude Opus; first-week calls jumped 27x

vs_Claude_Opus_4.8 ~parity, 80% cheaper
General reasoningCodingCost-cutting Claude replacement

How to access

Open weights on Hugging Face. API at bigmodel.cn and on OpenRouter.

MITCN
🇨🇳FlagshipOpen · Apache 2.0

Qwen3 Max

Alibaba

Alibaba's open-weight champion and the most-downloaded open model family on Earth. Strongest Chinese-language and agentic performance among open models.

36×cheaper than GPT-5.5 (output $/M)

Context

256K Tokens

Max Output

32K Tokens

Parameters

235B MoE (22B active)

Availability

Global

Output / 1M tokens$0.83USD
Input $0.28

※ Open-weight family with 1B–235B sizes; ~1B downloads (50%+ of global open-model downloads).

MMLU (Chinese) 88.1HumanEval 82.4
Self-hostingMultilingual appsAgents & tool use

How to access

Open weights on Hugging Face & ModelScope. Hosted API via Alibaba DashScope (global) or OpenRouter. Self-host freely.

Apache 2.0CN
🇨🇳FlagshipOpen · Modified MIT

Kimi K2.5

Moonshot A

Moonshot's long-context powerhouse with a 2M-token window — the go-to for ingesting entire codebases or document libraries in one call.

27×cheaper than GPT-5.5 (output $/M)

Context

2M Tokens

Max Output

128K Tokens

Parameters

1T MoE (32B active)

Availability

China + HK (OpenRouter global)

Output / 1M tokens$1.1USD
Input $1.1

※ Premium long-context specialist; 2M-token window.

Long-context recall industry-leading
Massive document analysisCodebase understandingAgentic search

How to access

API at platform.moonshot.cn (China). Globally accessible via OpenRouter.

Modified MITCN
🇨🇳BudgetOpen · Apache 2.0

MiniMax M2.5

MiniMax

MiniMax's efficient workhorse, consistently in the global top three by API call volume. Excellent price-to-performance for consumer-scale apps.

75×cheaper than GPT-5.5 (output $/M)

Context

200K Tokens

Max Output

64K Tokens

Parameters

MoE

Availability

Global

Output / 1M tokens$0.4USD
Input $0.1

※ Among the cheapest frontier-class models; top-3 global by call volume.

Global call volume Top 3
Roleplay & chatHigh-throughput agentsMultimodal

How to access

API at platform.minimaxi.com (international). Also on OpenRouter.

Apache 2.0CN
🌍FlagshipClosed

Claude Opus 5

Anthropic

Released July 24, 2026. Near-Fable-5 performance at half the cost. 2x+ Frontier-Bench vs Opus 4.8. Default model on Claude Max. Supports effort ladder and Fast mode (2.5x speed).

Context

1M Tokens

Max Output

128K

Parameters

Not disclosed

Availability

Global (excluding mainland China)

Output / 1M tokens$25USD
Input $5

※ Same price as Opus 4.8; Fast mode $10/$50; Batch 50% off; cache hits $0.50/MTok (90% savings)

DeepSWE v1.1 68.8%FrontierCode v1.1 53.4%
Complex agentic codingEnterprise code reviewLong-horizon autonomous tasks

How to access

anthropic.com / AWS Bedrock / Google Cloud / Microsoft Foundry. Not directly accessible in mainland China.

ProprietaryWEST

Why global developers are switching to Chinese AI models

Since 2025, Chinese LLMs have surged from 4.5% to over 46% of US enterprise tokens on OpenRouter. The driver is simple economics: models like DeepSeek V4 and GLM-5.2 deliver near-frontier quality at 60–90% lower cost than GPT-5.5 or Claude Opus. Companies like Lindy, Coinbase, DoorDash and Siemens have already made the switch.

How to choose the right Chinese LLM

For coding agents and long-context reasoning, DeepSeek V4 Pro leads. For self-hosting and multilingual apps, Qwen3 Max offers the strongest open-weight ecosystem. GLM-5.2 is the cheapest Claude replacement, Kimi K2.5 owns the 2M-token long-context niche, and MiniMax M2.5 / Doubao 2.0 win on raw throughput cost. Filter by origin, open weights, or price above.

How to access Chinese AI models from outside China

Most Chinese frontier models are available globally without a Chinese payment method. DeepSeek, Qwen, GLM and MiniMax all accept international cards or run on OpenRouter and Together.ai. Qwen, DeepSeek and GLM also ship open weights on Hugging Face for full self-hosting. Each card above lists exact access steps.