Skip to content

Chinese LLM API Pricing Research — 2026-07-26

Token-based API pricing for Chinese frontier models. All prices in USD per million tokens unless noted. Sourced live from official docs on July 26, 2026.


1. DeepSeek V4 Pro

Source: https://api-docs.deepseek.com/quick_start/pricing

Metric Price
Input (cache hit) $0.003625 / 1M tokens
Input (cache miss) $0.435 / 1M tokens
Output $0.87 / 1M tokens
Context window 1M tokens
Max output 384K tokens
Concurrency limit 500

Notes: - Supports thinking and non-thinking modes - Compatible with OpenAI and Anthropic API formats - Features: JSON output, tool calls, chat prefix completion (beta), FIM completion (beta, non-thinking only)


2. Kimi K3 (Moonshot AI)

Source: https://platform.kimi.com/docs/pricing/chat-k3

Metric Price (CNY) Approx. USD
Input (cache hit) ¥2.00 / 1M tokens ~$0.28
Input (cache miss) ¥20.00 / 1M tokens ~$2.80
Output ¥100.00 / 1M tokens ~$14.00
Context window 1,048,576 tokens (1M)

Notes: - Flagship model, always reasoning with configurable reasoning_effort (low/high/max, default max) - Supports automatic context caching, tool calls, JSON mode, structured output - Web search (web_search) is being upgraded, not recommended at this time - Model ID: kimi-k3


3. GLM-5.2 (Zhipu AI / 智谱)

Source: https://open.bigmodel.cn/pricing

Metric Price (CNY) Approx. USD
Input (cache hit) ¥2.00 / 1M tokens ~$0.28
Input (cache miss) ¥8.00 / 1M tokens ~$1.12
Output ¥28.00 / 1M tokens ~$3.92
Context window 1M tokens
Max output 128K tokens
Cache hit price Limited-time free (promo)

Notes: - Flagship open-weight model from Zhipu AI - Coding SOTA among open-weight models; Code Arena #1 globally - Performs between Claude Opus 4.7 and 4.8 in long-context benchmarks - 1M solid lossless context, specifically trained for long-context coding agents - Used by Hugging Face to defend against OpenAI's agent attack (commercial frontier models' guardrails blocked the task) - Supports thinking mode, streaming, function call, context cache, structured output, MCP, tool streaming - Model ID: glm-5.2


4. Grok 4.5 (xAI)

Source: https://docs.x.ai/docs/pricing

Metric Price (USD) Approx. CNY
Input (< 200K) $2.00 / 1M tokens ~¥14.30
Input (≥ 200K) $4.00 / 1M tokens ~¥28.60
Cached input (< 200K) $0.30 / 1M tokens ~¥2.15
Output (< 200K) $6.00 / 1M tokens ~¥42.90
Output (≥ 200K) $12.00 / 1M tokens ~¥85.80
Context window 500K tokens

Notes: - Long-context pricing: requests with ≥200K prompt tokens billed at 2x - Batch API: no discount listed for Grok 4.5 - Priority processing available at 2x standard rates - Model ID: grok-4.5


5. GPT-5.6 Luna (OpenAI)

Source: https://platform.openai.com/docs/pricing

Metric Price (USD) Approx. CNY
Input $1.00 / 1M tokens ~¥7.15
Cached input $0.10 / 1M tokens ~¥0.72
Cache writes $1.25 / 1M tokens ~¥8.94
Output $6.00 / 1M tokens ~¥42.90
Context window 256K

Notes: - Budget tier of GPT-5.6 family - Batch pricing: 50% off (input $0.50, output $3.00) - Flex processing at batch rates - Model ID: gpt-5.6-luna


6. GPT-5.6 Terra (OpenAI)

Source: https://platform.openai.com/docs/pricing

Metric Price (USD) Approx. CNY
Input $2.50 / 1M tokens ~¥17.88
Cached input $0.25 / 1M tokens ~¥1.79
Cache writes $3.125 / 1M tokens ~¥22.34
Output $15.00 / 1M tokens ~¥107.25
Context window 256K

Notes: - Mid-tier of GPT-5.6 family (between Sol and Luna) - Batch pricing: 50% off (input $1.25, output $7.50) - Flex processing at batch rates - Model ID: gpt-5.6-terra


7. Gemini 3.6 Flash (Google)

Source: https://cloud.google.com/vertex-ai/generative-ai/pricing

Metric Price (USD) Approx. CNY
Input (text, image, video, audio) $1.50 / 1M tokens ~¥10.73
Cached input $0.15 / 1M tokens ~¥1.07
Output (text, reasoning) $7.50 / 1M tokens ~¥53.63
Context window 1M+ tokens

Notes: - Multimodal input: text, image, video, audio all billed at the same input rate - Output price includes both response and reasoning tokens - Global pricing shown; non-global regions may be slightly higher - Model ID: gemini-3.6-flash


8. Muse Spark 1.1 (Meta)

Source: Meta Model API

Metric Price (USD) Approx. CNY
Input $1.25 / 1M tokens ~¥8.94
Cached input $0.15 – $1.00 / 1M tokens ~¥1.07 – ¥7.15
Output $4.25 / 1M tokens ~¥30.39
Free trial $20 credits for new accounts

Notes: - Multiple cache tiers available for discounted input rates - New developer accounts receive $20 in free testing credits - Model ID: muse-spark-1.1


9. Comparison Summary

All prices per 1M tokens (cache miss). USD conversion at $1 ≈ ¥7.15 (July 2026).

Model Output $ Output ¥ Input $ Input ¥
DeepSeek V4 Pro $0.87 ¥6.22 $0.435 ¥3.11
GLM-5.2 ~$3.92 ¥28.00 ~$1.12 ¥8.00
Muse Spark 1.1 $4.25 ~¥30.39 $1.25 ~¥8.94
GPT-5.6 Luna $6.00 ~¥42.90 $1.00 ~¥7.15
Grok 4.5 $6.00 ~¥42.90 $2.00 ~¥14.30
Gemini 3.6 Flash $7.50 ~¥53.63 $1.50 ~¥10.73
Kimi K3 ~$14.00 ¥100.00 ~$2.80 ¥20.00
GPT-5.6 Terra $15.00 ~¥107.25 $2.50 ~¥17.88

Price competitiveness: DeepSeek is the cheapest — 7.3x less than Kimi K3 on input, 16x on output. GLM-5.2 sits in the middle, ~2.5x DeepSeek's price but ~3.5x cheaper than Kimi K3.


10. Notes on CNY → USD Conversion

Grok 4.5 and DeepSeek publish USD pricing directly. Kimi K3 and GLM-5.2 pricing is listed in CNY. Approximate conversion used: $1 ≈ ¥7.15 (July 2026). Actual USD billing may vary based on each provider's payment processor.

DeepSeek publishes USD pricing directly.