Chinese LLM API Pricing Research — 2026-07-26
Token-based API pricing for Chinese frontier models. All prices in USD per million tokens unless noted. Sourced live from official docs on July 26, 2026.
1. DeepSeek V4 Pro
Source: https://api-docs.deepseek.com/quick_start/pricing
| Metric | Price |
|---|---|
| Input (cache hit) | $0.003625 / 1M tokens |
| Input (cache miss) | $0.435 / 1M tokens |
| Output | $0.87 / 1M tokens |
| Context window | 1M tokens |
| Max output | 384K tokens |
| Concurrency limit | 500 |
Notes: - Supports thinking and non-thinking modes - Compatible with OpenAI and Anthropic API formats - Features: JSON output, tool calls, chat prefix completion (beta), FIM completion (beta, non-thinking only)
2. Kimi K3 (Moonshot AI)
Source: https://platform.kimi.com/docs/pricing/chat-k3
| Metric | Price (CNY) | Approx. USD |
|---|---|---|
| Input (cache hit) | ¥2.00 / 1M tokens | ~$0.28 |
| Input (cache miss) | ¥20.00 / 1M tokens | ~$2.80 |
| Output | ¥100.00 / 1M tokens | ~$14.00 |
| Context window | 1,048,576 tokens (1M) |
Notes:
- Flagship model, always reasoning with configurable reasoning_effort (low/high/max, default max)
- Supports automatic context caching, tool calls, JSON mode, structured output
- Web search (web_search) is being upgraded, not recommended at this time
- Model ID: kimi-k3
3. GLM-5.2 (Zhipu AI / 智谱)
Source: https://open.bigmodel.cn/pricing
| Metric | Price (CNY) | Approx. USD |
|---|---|---|
| Input (cache hit) | ¥2.00 / 1M tokens | ~$0.28 |
| Input (cache miss) | ¥8.00 / 1M tokens | ~$1.12 |
| Output | ¥28.00 / 1M tokens | ~$3.92 |
| Context window | 1M tokens | |
| Max output | 128K tokens | |
| Cache hit price | Limited-time free (promo) |
Notes:
- Flagship open-weight model from Zhipu AI
- Coding SOTA among open-weight models; Code Arena #1 globally
- Performs between Claude Opus 4.7 and 4.8 in long-context benchmarks
- 1M solid lossless context, specifically trained for long-context coding agents
- Used by Hugging Face to defend against OpenAI's agent attack (commercial frontier models' guardrails blocked the task)
- Supports thinking mode, streaming, function call, context cache, structured output, MCP, tool streaming
- Model ID: glm-5.2
4. Grok 4.5 (xAI)
Source: https://docs.x.ai/docs/pricing
| Metric | Price (USD) | Approx. CNY |
|---|---|---|
| Input (< 200K) | $2.00 / 1M tokens | ~¥14.30 |
| Input (≥ 200K) | $4.00 / 1M tokens | ~¥28.60 |
| Cached input (< 200K) | $0.30 / 1M tokens | ~¥2.15 |
| Output (< 200K) | $6.00 / 1M tokens | ~¥42.90 |
| Output (≥ 200K) | $12.00 / 1M tokens | ~¥85.80 |
| Context window | 500K tokens |
Notes:
- Long-context pricing: requests with ≥200K prompt tokens billed at 2x
- Batch API: no discount listed for Grok 4.5
- Priority processing available at 2x standard rates
- Model ID: grok-4.5
5. GPT-5.6 Luna (OpenAI)
Source: https://platform.openai.com/docs/pricing
| Metric | Price (USD) | Approx. CNY |
|---|---|---|
| Input | $1.00 / 1M tokens | ~¥7.15 |
| Cached input | $0.10 / 1M tokens | ~¥0.72 |
| Cache writes | $1.25 / 1M tokens | ~¥8.94 |
| Output | $6.00 / 1M tokens | ~¥42.90 |
| Context window | 256K |
Notes:
- Budget tier of GPT-5.6 family
- Batch pricing: 50% off (input $0.50, output $3.00)
- Flex processing at batch rates
- Model ID: gpt-5.6-luna
6. GPT-5.6 Terra (OpenAI)
Source: https://platform.openai.com/docs/pricing
| Metric | Price (USD) | Approx. CNY |
|---|---|---|
| Input | $2.50 / 1M tokens | ~¥17.88 |
| Cached input | $0.25 / 1M tokens | ~¥1.79 |
| Cache writes | $3.125 / 1M tokens | ~¥22.34 |
| Output | $15.00 / 1M tokens | ~¥107.25 |
| Context window | 256K |
Notes:
- Mid-tier of GPT-5.6 family (between Sol and Luna)
- Batch pricing: 50% off (input $1.25, output $7.50)
- Flex processing at batch rates
- Model ID: gpt-5.6-terra
7. Gemini 3.6 Flash (Google)
Source: https://cloud.google.com/vertex-ai/generative-ai/pricing
| Metric | Price (USD) | Approx. CNY |
|---|---|---|
| Input (text, image, video, audio) | $1.50 / 1M tokens | ~¥10.73 |
| Cached input | $0.15 / 1M tokens | ~¥1.07 |
| Output (text, reasoning) | $7.50 / 1M tokens | ~¥53.63 |
| Context window | 1M+ tokens |
Notes:
- Multimodal input: text, image, video, audio all billed at the same input rate
- Output price includes both response and reasoning tokens
- Global pricing shown; non-global regions may be slightly higher
- Model ID: gemini-3.6-flash
8. Muse Spark 1.1 (Meta)
Source: Meta Model API
| Metric | Price (USD) | Approx. CNY |
|---|---|---|
| Input | $1.25 / 1M tokens | ~¥8.94 |
| Cached input | $0.15 – $1.00 / 1M tokens | ~¥1.07 – ¥7.15 |
| Output | $4.25 / 1M tokens | ~¥30.39 |
| Free trial | $20 credits for new accounts |
Notes:
- Multiple cache tiers available for discounted input rates
- New developer accounts receive $20 in free testing credits
- Model ID: muse-spark-1.1
9. Comparison Summary
All prices per 1M tokens (cache miss). USD conversion at $1 ≈ ¥7.15 (July 2026).
| Model | Output $ | Output ¥ | Input $ | Input ¥ |
|---|---|---|---|---|
| DeepSeek V4 Pro | $0.87 | ¥6.22 | $0.435 | ¥3.11 |
| GLM-5.2 | ~$3.92 | ¥28.00 | ~$1.12 | ¥8.00 |
| Muse Spark 1.1 | $4.25 | ~¥30.39 | $1.25 | ~¥8.94 |
| GPT-5.6 Luna | $6.00 | ~¥42.90 | $1.00 | ~¥7.15 |
| Grok 4.5 | $6.00 | ~¥42.90 | $2.00 | ~¥14.30 |
| Gemini 3.6 Flash | $7.50 | ~¥53.63 | $1.50 | ~¥10.73 |
| Kimi K3 | ~$14.00 | ¥100.00 | ~$2.80 | ¥20.00 |
| GPT-5.6 Terra | $15.00 | ~¥107.25 | $2.50 | ~¥17.88 |
Price competitiveness: DeepSeek is the cheapest — 7.3x less than Kimi K3 on input, 16x on output. GLM-5.2 sits in the middle, ~2.5x DeepSeek's price but ~3.5x cheaper than Kimi K3.
10. Notes on CNY → USD Conversion
Grok 4.5 and DeepSeek publish USD pricing directly. Kimi K3 and GLM-5.2 pricing is listed in CNY. Approximate conversion used: $1 ≈ ¥7.15 (July 2026). Actual USD billing may vary based on each provider's payment processor.
DeepSeek publishes USD pricing directly.