Skip to content

GLM-5.3-Flash Cost Index — Provider Ranking and Session-Level Cost vs DeepSeek V4 Flash

Research page: 2026-08-29. All prices are USD per 1M tokens from the cited pages. GLM-5.3-Flash (Z.ai / Zhipu, 320B-A18B, MIT, natively multimodal, 1M context) was released 2026-08-26 after a week as the anonymous "Ox-Alpha" model. A 50% launch discount runs until 2026-09-09 16:00 UTC, after which list prices apply.

TL;DR

  • The claim "GLM-5.3-Flash is the most cost efficient model at lower cost right now" is half right. On Artificial Analysis it is the cheapest per Intelligence Index task among frontier-tier models (57 IQ at $0.09/task, vs DeepSeek V4 Flash 52/$0.11, Qwen3.8-Flash-Next 56/$0.10, Gemini 3.7 Flash 56/$0.40). But AA ranks it only #16/111 on cost per task (2/4 cost units), and for cache-heavy coding agents — like our two real sessions — DeepSeek V4 Flash and Qwen3.8-Flash-Next are cheaper.
  • Unlike DeepSeek, there is no hidden cache-hit markup in the GLM ecosystem: every major host matches or undercuts Z.ai's $0.03 cached-input price. The index is therefore nearly flat across providers; the real levers are the promo, batch discounts, realized cache-hit rate, and verbosity.
  • Provider index (Z.ai official = 100, lower cheaper): promo until Sep 9 = 50, Fireworks batch = 49, Makora = 82–85, Fireworks serverless = 97, everyone else at list parity = 100.
  • Per-session example at the measured hit rates (S1 h=99.51%, S2 h=96.76%): GLM official costs $0.463 / $0.050 per session vs DeepSeek V4 Flash official $0.215 / $0.038 — DeepSeek is 2.15x cheaper in S1, 1.31x in S2.

1. Claim check

Intelligence (AA Intelligence Index v4.1.1) is kept as a separate column and is not blended into the cost index, so the two questions stay answerable independently: "cheapest way to run this workload" (index) vs "how smart is it" (IQ column).

Model AA IQ AA cost per task Notes
GLM-5.3-Flash 57 $0.09 Frontier-cheap on AA's task mix; verbose (150M tokens)
DeepSeek V4 Flash 0731 52 $0.11 (peak) AA snapshot uses peak rates; off-peak is cheaper
Qwen3.8-Flash-Next 56 $0.10 Same-day rival; 256K context vs GLM's 1M
Gemini 3.7 Flash 56 $0.40 Fast (324 t/s) and concise, but 4–7x the cost

GLM-5.3-Flash is the best intelligence-per-dollar of the frontier set, which is the kernel of truth in the claim. It is not the cheapest model overall — 15 cheaper models exist on AA's cost ranking, mostly tiny/old/free ones — and cache-heavy agent workloads flip the ranking below (Section 4).

2. Pricing (USD / 1M tokens, snapshot 2026-08-28)

Route Cache hit Fresh input Output
Z.ai official (list / promo to Sep 9) $0.03 / $0.015 $0.15 / $0.075 $0.50 / $0.25
Makora (cheapest OpenRouter endpoint) $0.024 $0.14 $0.47
Fireworks serverless (direct) $0.029 $0.15 $0.50
Fireworks batch (50% off serverless) $0.0145 $0.075 $0.25
DeepInfra / Novita / SiliconFlow / Together / DigitalOcean / Cloudflare / GMICloud / Baseten (OpenRouter endpoints) $0.03 $0.15 $0.50

Promo passthrough (until Sep 9): Z.ai, OpenRouter, DeepInfra, Novita, GMICloud, Relace. Everything doubles after the promo.

3. Provider cost index

Same formula as the DeepSeek page: E = r·[h·H + (1−h)·M] + O, index = provider / official × 100 (official = 100, lower = cheaper). Calibrated on the two real sessions:

Session Cached input Fresh input Output h r
S1 "add review layer" 13,521,536 67,079 95,292 99.51% 142.6
S2 "review repo vs requirements" 1,118,080 37,484 21,794 96.76% 53.0
Route S1 index S2 index
Promo until Sep 9 (official + passthrough) 50 50
Fireworks batch 48.5 48.9
Makora 81.7 84.5
Fireworks serverless 97.1 97.8
All other hosts at list parity 100 100

The DeepSeek pattern (third parties charging 1–4x official cache-hit prices) does not appear here: Z.ai's cached-input discount is only ~5x (vs DeepSeek's ~30x), so hosts have no room — or no incentive — to mark it up.

4. Cross-model comparison (IQ separate, not blended)

Same two sessions, official list prices. Index = our cost formula, GLM = 100.

Model AA IQ S1 index S2 index
GLM-5.3-Flash (Z.ai) 57 100 100
DeepSeek V4 Flash (official, 25% peak blend) 52 46.5 76.0
Qwen3.8-Flash-Next (QwenCloud intl) 56 60.0 68.5
Gemini 3.7 Flash (intro rate) 56 307 387
Kimi K2.6 (Moonshot) 563 603

At these hit rates the cost gap (1.7–2.2x) dwarfs the IQ gap (52–57, ≤10%), so DeepSeek and Qwen stay cheaper despite lower IQ. GLM wins when cache-hit ratio drops: crossover vs DeepSeek is h ≈ 87% at r=143, h ≈ 90% at r=53. Chat workloads (r≈4–10) favor GLM at any hit rate.

5. Cost per session, GLM-5.3-Flash vs DeepSeek V4 Flash

Actual dollar cost of the two real sessions (same token counts, route prices):

Route S1 S2 Δ vs GLM list (S1 / S2)
GLM-5.3-Flash official (list) $0.463 $0.050
GLM promo (until Sep 9) $0.232 $0.025 −$0.232 / −$0.025
GLM via Makora $0.379 $0.042 −$0.085 / −$0.008
DeepSeek V4 Flash official (25% peak blend) $0.215 $0.038 −$0.248 / −$0.012
DeepSeek V4 Flash official (all off-peak) $0.172 $0.030 −$0.291 / −$0.020
Qwen3.8-Flash-Next (QwenCloud) $0.278 $0.034 −$0.185 / −$0.016

So at the measured hit rates, running GLM-5.3-Flash at list instead of DeepSeek V4 Flash costs $0.248 more per S1-type session (2.15x) and $0.012 more per S2-type session (1.31x). The promo closes most of the S1 gap ($0.232 vs $0.215) but does not beat DeepSeek.

Hit-rate sensitivity (S1 token profile, GLM list vs DeepSeek 25% peak)

h GLM DeepSeek Cheaper by
99.5% $0.463 $0.215 DeepSeek $0.248
96.8% $0.508 $0.315 DeepSeek $0.193
90% $0.618 $0.559 DeepSeek $0.059
~87% $0.667 $0.668 crossover
85% $0.700 $0.740 GLM $0.040
70% $0.944 $1.283 GLM $0.338

The crossover sits at h ≈ 87–90% for agent-shaped workloads: below it, GLM's cheap fresh-input/output prices win; above it, DeepSeek's aggressive cache discount wins.

6. Which provider to use

  • Before Sep 9: Z.ai official or any promo passthrough (index 50). Lock in the discount before it doubles.
  • After Sep 9, GLM specifically: Fireworks direct — index ~97, zero-data retention, 67 t/s vs official's 49 t/s, and batch at ~49 for async jobs. DeepInfra at parity is the ZDR runner-up. Makora is the cheapest endpoint (index 82–85) but a small host — verify reliability and retention.
  • Pure cost on cache-heavy agents: DeepSeek V4 Flash official (index 46–76) beats every GLM route. Choose GLM for IQ 57, native multimodal input, or 1M context.
  • OpenRouter aggregates to the cheapest endpoint (Makora post-promo) with no per-token markup, but its load-balancing erodes realized cache-hit rates — pin one provider for cache-heavy workloads.

7. Research note: Vercel AI Gateway for cache-light tasks

For cache-light one-shot work (h ≈ 0, e.g. research summaries), the practical strategy is two routes: an aggregator for one-shot tasks and official DeepSeek V4 Flash for cache-heavy coding agents. Vercel AI Gateway is a reasonable zero-cost middle layer for the first half (verified 2026-08-29).

Cost. Zero markup and zero platform fee on tokens, including BYOK, on all plans. GLM-5.3-Flash is served at pass-through prices — list $0.15 / $0.50 / $0.03, promo ~$0.08 / $0.25 / $0.02 — with one upstream host even cheaper ($0.10 / $0.40 / $0.01) and one premium ($0.45 / $1.50 / $0.09). "Free" is a funnel, not charity: Vercel monetizes the platform (hosting, prepaid credit float, paid compliance controls), not tokens.

Routing is dynamic by default. AI Gateway picks upstream providers by recent uptime and latency, not price, and can change per request — the same model can cost $0.10–$0.45 input depending on the host. Pin a provider per request (provider slug, or the order / only / sort options) when cost, speed, or cache continuity matters.

Credentials. No Z.ai key needed: Vercel's system credentials call the provider and bill you at API rates through one Vercel key, no markup. BYOK (paid tier) is optional — for enterprise agreements, existing credits, or ZDR under your own account.

Data. Vercel does not train on prompts, and AI Gateway uses zero data retention by default (prompts/responses deleted after completion). Per-request zeroDataRetention and disallowPromptTraining are free; team-wide ZDR enforcement costs $0.10/M and routes only to providers with verified ZDR agreements. Usage telemetry (model, tokens, latency, provider) is still kept for dashboards.

Requirements / caveats. Free tier: $5/month credits (the credit stops permanently on first top-up, then PAYG) and lower per-model rate limits. BYOK is paid-tier only; credits are prepaid float; the gateway adds an infrastructure dependency and usage visibility.

Coding route. Official DeepSeek V4 Flash stays the simplest and cheapest for cache-heavy agents (index 46–76 at the calibrated sessions). Watch the weekday peak-hour premium (2x) and China-hosted retention; Fireworks is the flat-price ZDR alternative.

8. Caveats

  • Promo ends 2026-09-09 16:00 UTC; cached-input storage free is limited-time.
  • reasoning_effort defaults to max (reasoning bills as output); AA measured GLM as verbose (150M tokens vs 110M median) — set low where possible.
  • Official Z.ai API is slow (49 t/s); DigitalOcean (88 t/s), Makora (69), Fireworks (67) are faster on OpenRouter.
  • Z.ai is a Chinese-first provider with no published retention window; Western hosts (Fireworks, DeepInfra, Together, DigitalOcean, Cloudflare) offer ZDR.
  • Watch Qwen3.8-Flash-Next: cheaper cache ($0.0165) and output ($0.47), near-equal IQ (56), but 256K context and a restricted open-weight license.

9. Sources

  • Z.ai pricing docs (docs.z.ai/guides/overview/pricing, snapshot 2026-08-27)
  • OpenRouter endpoints API for z-ai/glm-5.3-flash-20260826 (snapshot 2026-08-28)
  • DeepInfra model page (promo + list), Fireworks model page + Batch API docs, Cloudflare Workers AI model docs
  • Artificial Analysis model pages: GLM-5.3-Flash, DeepSeek V4 Flash 0731, Qwen3.8-Flash-Next, Gemini 3.7 Flash (Aug 2026)
  • DeepSeek official pricing (peak/off-peak, 2026-08-23 snapshot; see companion page)
  • Qwen3.8-Flash pricing: QwenCloud (runtimewire, jiemian 2026-08-26/27)
  • Kimi K2.6 pricing page (Moonshot); Apidog GLM-5.3-Flash pricing analysis; Pandaily "Flash models redraw China's LLM flagship line" (2026-08-27)
  • Vercel AI Gateway: pricing, security-and-compliance, ZDR and disallow-prompt- training docs, provider options, GLM 5.3 Flash model page + changelog (2026-08-25/29)