Beyond Anthropic, OpenAI, and Google — The OpenRouter Provider Field
Research date: 2026-09-05. All token figures are the OpenRouter providers page snapshot as fetched Sep 5, 2026, sorted by daily tokens. The big-three company routes — OpenAI, Google Vertex/Google AI Studio, Anthropic/Claude Platform on AWS — are intentionally removed for this post. For the big-three comparison, see Which AI API Provider Is Most Popular for Coding-Agent Inference?.
TL;DR
- After Anthropic, OpenAI and Google, the largest OpenRouter provider routes by daily tokens are Tencent Cloud (3.7T/day), GMICloud (1.3T/day), Z.ai (982.6B/day), NVIDIA (723.6B/day) and NovitaAI (654.3B/day).
- The field is not one market. It splits into four groups: official model-lab APIs (Tencent, Z.ai, DeepSeek, Baidu, Alibaba), serverless GPU/inference hosts (GMICloud, Novita, DeepInfra, and below them Together, Fireworks, Baseten, Modal, Parasail), cloud marketplaces (Amazon Bedrock, Azure), and coding-native specialists (Relace, NVIDIA NIM, StreamLake/KAT-Coder).
- Token volume is not revenue. Several of these routes exist because their tokens are cheap; DeepInfra, for example, has the largest OpenRouter catalog (103 models) and 16.1T monthly tokens, while official DeepSeek processes 11.5T/month. On spend, Anthropic still dwarfs this entire table (see the companion popularity post).
- If you care about proprietary code, the policy flags matter more than volume: most of the big non-big-three routes have zero retention and do not train on prompts, but DeepSeek and NVIDIA are flagged as training, and several Chinese official APIs retain prompts.
1. The ranking (OpenRouter, Sep 5, 2026)
OpenRouter's providers page lists 83 provider routes with live token volume, policy flags, BYOK support, and model counts. This table keeps the twelve biggest routes after dropping OpenAI, Google Vertex, Google AI Studio, Anthropic and Claude Platform on AWS.
| Provider route | HQ | Daily tokens | Monthly tokens | Models | Trains? | Retention | BYOK |
|---|---|---|---|---|---|---|---|
| Tencent Cloud | China | 3.74T | 44.28T | 5 | No | Zero retention | Yes |
| GMICloud | United States | 1.26T | 13.82T | 23 | No | Retains prompts | Yes |
| Z.ai | Singapore | 982.6B | 11.98T | 13 | No | Zero retention | Yes |
| NVIDIA | United States | 723.6B | 21.14T | 8 | Yes | Retains prompts | Yes |
| NovitaAI | United States | 654.3B | 19.57T | 72 | No | Zero retention | Yes |
| StreamLake | China | 617.9B | 11.02T | 22 | No | Retains prompts | Yes |
| DeepSeek | China | 516.0B | 11.52T | 3 | Yes | Retains prompts | Yes |
| Relace | — (SF-based startup) | 438.5B | 6.35T | 3 | No | Zero retention | Yes |
| Amazon Bedrock | United States | 427.0B | 9.69T | 32 | No | Zero retention | Yes |
| Baidu Qianfan | China | 420.2B | 11.26T | 9 | No | Retains prompts | Yes |
| Alibaba Cloud Int. | Singapore | 416.8B | 8.27T | 62 | No | Retains prompts | Yes |
| DeepInfra | United States | 346.6B | 16.09T | 103 | No | Zero retention | Yes |
How to read this:
- These are OpenRouter-routed tokens, not each provider's total direct API traffic. A lab such as DeepSeek sells far more tokens through its own API than it moves through OpenRouter.
- "Provider" here means the billing/endpoint route. The same model often appears on several routes: DeepSeek V4 Flash 0731, GLM-5.3-Flash and MiMo-V2.5-Pro are served by DeepSeek/Z.ai/Xiaomi directly and by hosts like GMICloud, StreamLake, Novita and DeepInfra.
- Monthly and daily rankings disagree (e.g., NVIDIA is 4th by daily tokens but 2nd by monthly; DeepInfra is 12th by daily but 4th by monthly), so do not treat one window as the full story.
- "Trains?" and retention flags come from OpenRouter's structured view of each provider's own terms. They say nothing about what OpenRouter itself logs or what your agent harness stores.
2. The four markets hiding in this table
2.1 Official model labs selling their own models
These are the model owners, not resellers. They dominate when an open-weight or frontier model's own API is competitive enough to win routing:
| Provider | What it is | Notes |
|---|---|---|
| Tencent Cloud | Tencent's international cloud (Hunyuan) | Five text models including Hunyuan Hy3/Hy4-generation lines; largest daily and monthly token volume outside the big three; zero retention, no training, BYOK |
| Z.ai | Singapore-based international brand of Zhipu AI (GLM family) | 13 models including GLM-5.3-Flash (the former "Ox Alpha"), GLM-5.2/5.3, and multimodal GLM; zero retention, no training, BYOK |
| DeepSeek | Official DeepSeek API | Only three model listings (the V4 family and one multimodal entry) but huge volume per model; OpenRouter flags it as training on prompts and retaining them |
| Baidu Qianfan | Baidu AI Cloud model platform | Nine models, primarily the ERNIE/Qianfan catalog; retains prompts, no training |
| Alibaba Cloud Int. | Alibaba Cloud's international arm (Qwen) | 62 models, one of the deepest text+multimodal catalogs among labs; retains prompts, no training |
Lower down the same list: Xiaomi (only two models but 29.6T monthly tokens — behind only Tencent in this whole field), MiniMax (13 models, 5.5T/month), and Moonshot AI (Kimi family, 2.5T/month).
The coding relevance: DeepSeek, Z.ai/GLM and Tencent/Hunyuan have been the most adopted Chinese model families in OpenRouter agent traffic through 2026. They are also the ones whose official APIs are often beaten on price by the hosts below.
2.2 Serverless GPU and inference hosts
These providers do not own the frontier models; they serve open weights (and sometimes white-label models) cheaply and with developer-friendly APIs:
| Provider | What it is | Notes |
|---|---|---|
| GMICloud | US GPU-cloud/inference platform (Mountain View, CA) | Reports ~$500M+ contracted ARR in 2026 and ~2.5T tokens/week in its inference business; serves DeepSeek, GLM, MiniMax, Moonshot and other open weights across US/APAC endpoints; OpenRouter flags it as retaining prompts |
| NovitaAI | San Francisco serverless inference + GPU cloud + agent sandboxes | Bootstrapped, founded late 2023; 72 models on OpenRouter out of a larger 200+ model platform; official Hugging Face Inference Partner; added a Firecracker-based Agent Sandbox in Apr 2026 |
| DeepInfra | Palo Alto serverless inference platform | Largest model count in this table (103 routes across text, image, video, audio, embeddings); zero retention, no training, BYOK; already reviewed in the DeepSeek V4 Flash provider cost index |
The rest of the same group, further down the OpenRouter page: Parasail (38 models), Together (33 models), SiliconFlow (41 models, Singapore route), Fireworks (13 models), Baseten (13 models), Modal (4 models), CoreWeave (17 models), and DigitalOcean (19 models). Groq, Cerebras and SambaNova are the speed-specialist cousins with much smaller OpenRouter volume.
This group is where the "one key, many models" workflow comes from. For agent coding specifically, it is also where you can get a cheaper DeepSeek or GLM route than the lab itself charges — which is exactly why these hosts capture so many tokens at so little spend.
2.3 Coding-native and task specialists
The surprise in the table is Relace: 438.5B daily tokens from only three models. Relace (officially Squack Inc.) is a San Francisco coding-agent infrastructure startup founded by Preston Zhou and Eitan Borgnia, with a $23M Series A led by a16z (SiliconANGLE, YC launch). It does not sell a frontier general-purpose LLM. Its models are specialized helper models for coding agents — codebase search/retrieval, reranking, and an "apply" model that merges AI edits into files at thousands of tokens per second. That is a strong signal that OpenRouter's coding-agent traffic increasingly consumes non-frontier utility models, not just flagship reasoning models.
NVIDIA NIM is the second coding-adjacent route: eight models (including its own Nemotron lines) with free tiers, but OpenRouter flags NVIDIA's endpoint as training on prompts and retaining them, so enterprise coding teams usually route around it.
StreamLake is the third: the tech-commercialization cloud brand of Kuaishou (Kwai), not ByteDance. Kuaishou's StreamLake launched an AI-coding product matrix in Oct 2025 — CodeFlicker, the KAT-Coder model family, and a model platform — and its KAT-Coder models are among the coding models available through this OpenRouter route (Kuaishou StreamLake launch coverage, KAT-Coder Pro v2.5 pricing). Its 22 models and 617.9B/day include both Kuaishou models and hosted third-party open weights.
2.4 Cloud marketplaces
Amazon Bedrock (32 models) and, further down, Azure (57 models) are enterprise marketplaces that route many labs' models through one cloud contract. They matter for API-key agent inference because they are how many companies consume Claude, GPT, or Qwen without a direct vendor relationship. Do not confuse them with "non-big-three" models: a Bedrock route can serve Anthropic or OpenAI models, so marketplace routes are not a guarantee that you have left the big-three model families.
3. What the policy flags actually mean
OpenRouter's provider page reports two policy dimensions:
- Trains — whether the provider says it may train on your prompts. In this table, only DeepSeek and NVIDIA are flagged yes. OpenRouter lets users set account-level filters that refuse routing to such providers.
- Retention — what the provider says it does with prompts after the call: zero retention, retain for 30/55 days, or retain indefinitely. Most US serverless hosts advertise zero retention; several Chinese official APIs retain prompts; GMICloud is the notable US host with a "retains prompts" flag in OpenRouter's view.
Three cautions:
- These flags reflect each provider's stated policy as OpenRouter reads it, not an audit.
- BYOK does not bypass OpenRouter routing entirely; the request still passes through OpenRouter unless you call the provider directly.
- For agent coding, the sensitive content is source code, repo structure and terminal output. A "zero retention" provider can still be a bad fit if your compliance rule is about jurisdiction, not retention. Compare HQ and data regions separately from the retention flag.
4. Bottom line
If Anthropic, OpenAI and Google are excluded, the OpenRouter field's biggest providers are:
- By daily volume: Tencent Cloud, GMICloud, Z.ai, NVIDIA, NovitaAI.
- By monthly volume: Tencent Cloud, NVIDIA, NovitaAI, DeepInfra, GMICloud.
- By model breadth: DeepInfra (103), Novita (72), Alibaba Cloud (62), Azure (57), SiliconFlow (41), Parasail (38).
- By coding specialization: Relace and StreamLake/KAT-Coder are the purest "agent-coding" providers; the rest win by serving coding-capable open-weight models (DeepSeek, GLM, Qwen, Hunyuan, MiMo) cheaply.
For a practical agent stack, the most useful reading is not "which one is biggest" but which category you need: official lab API when you want the model-maker's own SLA and feature set, a serverless host when you want the same open-weight model cheaper or with zero retention, a marketplace when you want one enterprise contract, or a specialist like Relace when you want to add retrieval/apply infrastructure around a frontier model rather than replace it.
Sources
- OpenRouter providers page (live snapshot Sep 5, 2026): openrouter.ai/providers
- OpenRouter provider data-retention/logging documentation: openrouter.ai/docs/guides/privacy/logging
- GMI Cloud ARR/inference growth reporting: Daily AI Brief, Data Center Dynamics
- Novita profile and April 2026 launches: Benched.ai, Malaysian Reserve
- Relace company and products: SiliconANGLE, YC launch, relace.ai
- Kuaishou StreamLake AI-coding matrix and KAT-Coder: CCIDnet, TokenCost, TheBlockBeats
- Z.ai/Zhipu background: Business Insider
- DeepInfra profile: DataLearner
- Companion: Which AI API Provider Is Most Popular for Coding-Agent Inference?