Chinese LLM API Pricing Research — 2026-07-26

Token-based API pricing for Chinese frontier models, in USD per million tokens, sourced live from official docs on July 26, 2026.

Token-based API pricing for Chinese frontier models. All prices in USD per million tokens unless noted. Sourced live from official docs on July 26, 2026.


1. DeepSeek V4 Pro

Source: https://api-docs.deepseek.com/quick_start/pricing

MetricPrice
Input (cache hit)$0.003625 / 1M tokens
Input (cache miss)$0.435 / 1M tokens
Output$0.87 / 1M tokens
Context window1M tokens
Max output384K tokens
Concurrency limit500

Notes:

  • Supports thinking and non-thinking modes
  • Compatible with OpenAI and Anthropic API formats
  • Features: JSON output, tool calls, chat prefix completion (beta), FIM completion (beta, non-thinking only)

2. Kimi K3 (Moonshot AI)

Source: https://platform.kimi.com/docs/pricing/chat-k3

MetricPrice (CNY)Approx. USD
Input (cache hit)¥2.00 / 1M tokens~$0.28
Input (cache miss)¥20.00 / 1M tokens~$2.80
Output¥100.00 / 1M tokens~$14.00
Context window1,048,576 tokens (1M)

Notes:

  • Flagship model, always reasoning with configurable reasoning_effort (low/high/max, default max)
  • Supports automatic context caching, tool calls, JSON mode, structured output
  • Web search (web_search) is being upgraded, not recommended at this time
  • Model ID: kimi-k3

3. GLM-5.2 (Zhipu AI / 智谱)

Source: https://open.bigmodel.cn/pricing

MetricPrice (CNY)Approx. USD
Input (cache hit)¥2.00 / 1M tokens~$0.28
Input (cache miss)¥8.00 / 1M tokens~$1.12
Output¥28.00 / 1M tokens~$3.92
Context window1M tokens
Max output128K tokens
Cache hit priceLimited-time free (promo)

Notes:

  • Flagship open-weight model from Zhipu AI
  • Coding SOTA among open-weight models; Code Arena #1 globally
  • Performs between Claude Opus 4.7 and 4.8 in long-context benchmarks
  • 1M solid lossless context, specifically trained for long-context coding agents
  • Used by Hugging Face to defend against OpenAI’s agent attack (commercial frontier models’ guardrails blocked the task)
  • Supports thinking mode, streaming, function call, context cache, structured output, MCP, tool streaming
  • Model ID: glm-5.2

4. Grok 4.5 (xAI)

Source: https://docs.x.ai/docs/pricing

MetricPrice (USD)Approx. CNY
Input (< 200K)$2.00 / 1M tokens~¥14.30
Input (≥ 200K)$4.00 / 1M tokens~¥28.60
Cached input (< 200K)$0.30 / 1M tokens~¥2.15
Output (< 200K)$6.00 / 1M tokens~¥42.90
Output (≥ 200K)$12.00 / 1M tokens~¥85.80
Context window500K tokens

Notes:

  • Long-context pricing: requests with ≥200K prompt tokens billed at 2x
  • Batch API: no discount listed for Grok 4.5
  • Priority processing available at 2x standard rates
  • Model ID: grok-4.5

5. GPT-5.6 Luna (OpenAI)

Source: https://platform.openai.com/docs/pricing

MetricPrice (USD)Approx. CNY
Input$1.00 / 1M tokens~¥7.15
Cached input$0.10 / 1M tokens~¥0.72
Cache writes$1.25 / 1M tokens~¥8.94
Output$6.00 / 1M tokens~¥42.90
Context window256K

Notes:

  • Budget tier of GPT-5.6 family
  • Batch pricing: 50% off (input $0.50, output $3.00)
  • Flex processing at batch rates
  • Model ID: gpt-5.6-luna

6. GPT-5.6 Terra (OpenAI)

Source: https://platform.openai.com/docs/pricing

MetricPrice (USD)Approx. CNY
Input$2.50 / 1M tokens~¥17.88
Cached input$0.25 / 1M tokens~¥1.79
Cache writes$3.125 / 1M tokens~¥22.34
Output$15.00 / 1M tokens~¥107.25
Context window256K

Notes:

  • Mid-tier of GPT-5.6 family (between Sol and Luna)
  • Batch pricing: 50% off (input $1.25, output $7.50)
  • Flex processing at batch rates
  • Model ID: gpt-5.6-terra

7. Gemini 3.6 Flash (Google)

Source: https://cloud.google.com/vertex-ai/generative-ai/pricing

MetricPrice (USD)Approx. CNY
Input (text, image, video, audio)$1.50 / 1M tokens~¥10.73
Cached input$0.15 / 1M tokens~¥1.07
Output (text, reasoning)$7.50 / 1M tokens~¥53.63
Context window1M+ tokens

Notes:

  • Multimodal input: text, image, video, audio all billed at the same input rate
  • Output price includes both response and reasoning tokens
  • Global pricing shown; non-global regions may be slightly higher
  • Model ID: gemini-3.6-flash

8. Muse Spark 1.1 (Meta)

Source: Meta Model API

MetricPrice (USD)Approx. CNY
Input$1.25 / 1M tokens~¥8.94
Cached input$0.15 – $1.00 / 1M tokens~¥1.07 – ¥7.15
Output$4.25 / 1M tokens~¥30.39
Free trial$20 credits for new accounts

Notes:

  • Multiple cache tiers available for discounted input rates
  • New developer accounts receive $20 in free testing credits
  • Model ID: muse-spark-1.1

9. Comparison Summary

All prices per 1M tokens (cache miss). USD conversion at $1 ≈ ¥7.15 (July 2026).

ModelOutput $Output ¥Input $Input ¥
DeepSeek V4 Pro$0.87¥6.22$0.435¥3.11
GLM-5.2~$3.92¥28.00~$1.12¥8.00
Muse Spark 1.1$4.25~¥30.39$1.25~¥8.94
GPT-5.6 Luna$6.00~¥42.90$1.00~¥7.15
Grok 4.5$6.00~¥42.90$2.00~¥14.30
Gemini 3.6 Flash$7.50~¥53.63$1.50~¥10.73
Kimi K3~$14.00¥100.00~$2.80¥20.00
GPT-5.6 Terra$15.00~¥107.25$2.50~¥17.88

Price competitiveness: DeepSeek is the cheapest — 7.3x less than Kimi K3 on input, 16x on output. GLM-5.2 sits in the middle, ~2.5x DeepSeek’s price but ~3.5x cheaper than Kimi K3.


10. Notes on CNY → USD Conversion

Grok 4.5 and DeepSeek publish USD pricing directly. Kimi K3 and GLM-5.2 pricing is listed in CNY. Approximate conversion used: $1 ≈ ¥7.15 (July 2026). Actual USD billing may vary based on each provider’s payment processor.

DeepSeek publishes USD pricing directly.