Chinese LLM API Pricing Research — 2026-09-12
Token-based API pricing for Chinese frontier models, in USD per million tokens, updated from official docs on September 12, 2026.
Token-based API pricing for Chinese frontier models. All prices in USD per million tokens unless noted. Updated from official docs on September 12, 2026.
Unified pricing comparison
All prices are per 1M tokens. USD is shown first; CNY equivalents use $1 ≈ ¥7.15 for comparison. “Uncached input” means cache miss. A dash means that the source did not list a separate rate.
| Model | Uncached input | Cached input | Cache write | Output | Context and limits | Source |
|---|---|---|---|---|---|---|
| DeepSeek V4 Pro | $0.435 (¥3.11) | $0.003625 | — | $0.87 (¥6.22) | 1M; max output 384K; concurrency 500 | official pricing |
| Hy4 preview (Tencent Hunyuan) | $0.834 (≈¥5.96) | $0.042 (≈¥0.30) | — | $2.501 (≈¥17.88) | 1,048,576; max output 64K | official announcement and pricing |
| GLM-5.2 (Zhipu AI) | $1.12 (¥8.00) | $0.28 (¥2.00); limited-time free promo | — | $3.92 (¥28.00) | 1M; max output 128K | official pricing |
| Muse Spark 1.1 (Meta) | $1.25 (≈¥8.94) | $0.15–$1.00 (≈¥1.07–¥7.15) | — | $4.25 (≈¥30.39) | $20 credits for new accounts | Meta Model API |
| GPT-5.6 Luna (OpenAI) | $1.00 (≈¥7.15) | $0.10 (≈¥0.72) | $1.25 (≈¥8.94) | $6.00 (≈¥42.90) | 256K | official pricing |
| Grok 4.5 (xAI) | $2.00 (<200K); $4.00 (≥200K) | $0.30 (<200K) | — | $6.00 (<200K); $12.00 (≥200K) | 500K; ≥200K prompts billed at 2x | official pricing |
| Gemini 3.6 Flash (Google) | $1.50 (≈¥10.73) | $0.15 (≈¥1.07) | — | $7.50 (≈¥53.63) | 1M+; text, image, video, audio input | Vertex AI pricing |
| Kimi K3 (Moonshot AI) | $2.80 (¥20.00) | $0.28 (¥2.00) | — | $14.00 (¥100.00) | 1,048,576 (1M) | official pricing |
| GPT-5.6 Terra (OpenAI) | $2.50 (≈¥17.88) | $0.25 (≈¥1.79) | $3.125 (≈¥22.34) | $15.00 (≈¥107.25) | 256K | official pricing |
Price competitiveness: DeepSeek remains the cheapest on both uncached input and output. Hy4 preview is next at $0.834 input and $2.501 output, followed by GLM-5.2 at $1.12 and $3.92. Kimi K3 is about 6.4× DeepSeek’s price on uncached input and 16.1× on output.
Model notes
DeepSeek V4 Pro
- Supports thinking and non-thinking modes.
- Compatible with OpenAI and Anthropic API formats.
- Features JSON output, tool calls, chat prefix completion (beta), and FIM completion (beta, non-thinking only).
Hy4 preview
- Tencent’s early Hunyuan 4 preview is available through Tencent Cloud and OpenRouter.
- Model ID:
hy4-preview.
Kimi K3
- Flagship model, always reasoning with configurable
reasoning_effort(low/high/max, default max). - Supports automatic context caching, tool calls, JSON mode, and structured output.
- Web search (
web_search) is being upgraded and is not recommended at this time. - Model ID:
kimi-k3.
GLM-5.2
- Flagship open-weight model from Zhipu AI.
- Supports thinking mode, streaming, function call, context cache, structured output, MCP, and tool streaming.
- Model ID:
glm-5.2.
Grok 4.5
- Batch API discount was not listed for Grok 4.5.
- Priority processing is available at 2× standard rates.
- Model ID:
grok-4.5.
GPT-5.6 Luna and Terra
- Batch pricing is 50% off for both models; Flex processing is available at batch rates.
- Luna is the budget tier; Terra is the mid-tier of the GPT-5.6 family.
- Model IDs:
gpt-5.6-lunaandgpt-5.6-terra.
Gemini 3.6 Flash
- Multimodal input—text, image, video, and audio—is billed at the same input rate.
- Output pricing includes both response and reasoning tokens.
- Global pricing is shown; non-global regions may be slightly higher.
- Model ID:
gemini-3.6-flash.
Muse Spark 1.1
- Multiple cache tiers are available for discounted input rates.
- New developer accounts receive $20 in free testing credits.
- Model ID:
muse-spark-1.1.
Notes on CNY → USD conversion
Kimi K3 and GLM-5.2 list pricing in CNY; the other models in this comparison publish USD pricing or are represented using their USD list price. Approximate conversion used: $1 ≈ ¥7.15. Actual USD billing may vary based on each provider’s payment processor.