Skip to content

Which AI API Provider Is Most Popular for Coding-Agent Inference?

Research date: 2026-09-05. All figures are snapshots, not audited totals. Provider popularity depends on whether you rank by routed token volume, dollar spend, developer installs, or paid coding-agent adoption — and each dataset sees a different slice of the market. Related context: The LLM ARR Debate (Sept 2026) covers Claude Code vs Codex revenue; DeepSeek V4 Flash provider cost index and GLM-5.3-Flash cost index cover the price side.

TL;DR

  • Anthropic (Claude) is the most popular provider for paid, quality-critical coding-agent inference. On the Vercel AI Gateway it holds 61% of June 2026 spend on 32% of tokens, ≥72% of spend in coding agents, back-office agents and app generation, and Claude Code is the largest paid coding product by tracked ARR (~\(15.1B vs Codex's ~\)8.8B as of Aug 10).
  • DeepSeek is the most popular provider by raw token traffic on neutral developer routers. The Sep 4 OpenRouter snapshot puts DeepSeek at 19.0% of tracked provider tokens, ahead of OpenAI (15.7%), Z.ai (15.0%) and Tencent (12.9%); Anthropic is 4.4%. The coding-agent slice on Vercel is even more extreme: DeepSeek ran 49% of coding-agent tokens in May on 4% of the cost.
  • OpenAI has the largest developer/API ecosystem by installs and reach — ~427M monthly openai PyPI downloads in Aug 2026 and the largest first-party direct-API disclosed scale — and its Codex/GPT-5.6 family is the fastest climber on intensity metrics. Google leads by first-party consumer volume but is nearly absent from Vercel coding-agent tokens (<2% in June).
  • The honest answer: there is no single winner unless you fix the metric. If "most popular" means which provider gets paid for coding/agent inference, it is Anthropic. If it means which provider's API keys move the most tokens through community routers, it is DeepSeek (with Z.ai/GLM, Tencent and Xiaomi close behind). If it means which provider developers install and build with most, it is still OpenAI.

1. The ranking question

"Which provider is most popular for AI API-key, LLM-agent, coding inference" is actually several questions:

Metric What it measures Main blind spot
Routed token volume (OpenRouter) What models developer traffic actually calls when switching is easy Skews toward cheap/high-volume models; misses direct APIs and first-party apps
Gateway token & spend share (Vercel AI Gateway) Production apps and agents that bring their own API keys Only the Vercel gateway's slice; spend normalized to list price
SDK downloads (PyPI/npm) What developers install/build with Install churn and OpenAI-compatible interfaces inflate counts
Enterprise spend/adoption surveys (Menlo, Ramp) Who enterprises pay and how many pay US-biased, survey/payment-card based, less current for model mix
Coding product ARR (Claude Code, Codex) Who gets paid for coding agents specifically Tracked estimates, not audited revenue

None of these is a census. Together they triangulate the answer.


2. Provider token traffic on OpenRouter (community routing)

OpenRouter is the largest neutral model router and publishes per-model rankings under CC BY 4.0 (openrouter.ai/rankings). It does not publish an official provider roll-up, so the table below is from Whatstrending, an independent tracker that aggregates OpenRouter's public top-model data by provider. Snapshot: Sep 4, 2026, covering 105.06T tracked tokens in a seven-day window.

Rank Provider 7-day tokens Share of tracked volume
1 DeepSeek 19.92T 19.0%
2 OpenAI 16.49T 15.7%
3 Z.ai 15.73T 15.0%
4 Tencent 13.57T 12.9%
5 Google 7.1T 6.8%
6 MiniMax 6.54T 6.2%
7 Xiaomi 6.43T 6.1%
8 NVIDIA 5.78T 5.5%
9 Anthropic 4.63T 4.4%
10 Moonshot 2.01T 1.9%
11 Poolside 1.53T 1.5%
12 Alibaba / Qwen 1.32T 1.3%
13 Others 4.01T 3.8%

How to read this:

  • The roll-up is by model family/provider, not by the infrastructure company serving the call. A developer using a DeepSeek model through Fireworks or Baseten counts as DeepSeek model tokens even though the inference bill goes to an American host.
  • OpenRouter traffic excludes ChatGPT, Claude.ai, Gemini and direct vendor API calls, so the large first-party reach of OpenAI, Google and Anthropic is undercounted here.
  • Token share is not spend share or quality share. Anthropic at #9 by tokens can still be #1 by dollars.

The OpenRouter model list shows the same pattern in fast motion. Chinese media recaps of the weekly OpenRouter leaderboard report:

  • Week ending Aug 30: weekly volume crossed 113T tokens, up ~17x from ~6.4T at the start of 2026; 13 of the top 20 models were Chinese, carrying ~75.4% of top-20 tokens (recap via Eastmoney's Future Tech).
  • Top models that week: Ox Alpha / GLM-5.3-Flash 15.70T, DeepSeek V4 Flash 0731 at 12.30T, Xiaomi MiMo-V2.5 at 9.14T, GPT-5.6 Luna at 7.79T, Tencent Hy3 at 6.66T.
  • Anthropic's Claude models have slid in OpenRouter top-20 token rankings: four Claude models in the top 20 a month earlier, two by Aug 30, at <3% of top-20 tokens combined — while Claude Code still leads the paid/agent application side.

Coding and agent tools dominate this platform. Agent-style apps take most of the OpenRouter app-leaderboard volume: the 30-day snapshot of Aug 24 ranks Hermes Agent at 41.29T tokens, Claude Code at 10.09T, Kilo Code at 8.08T, Cline at 5.30T, and Codex at 122.6B (Codesota OpenRouter app leaderboard). The Claude Code and Codex numbers are only their OpenRouter-routed slices; the products mostly run on direct Anthropic/OpenAI APIs, which are invisible to this dataset.


3. Vercel AI Gateway: token share vs spend share

Vercel's AI Gateway Production Index is a useful counter-sample because it sees production apps and agents that bring their own API keys (tens of trillions of routed tokens per month) rather than only router-hopping developers.

May 2026 (Vercel June 2026 index):

Provider Token share Spend share
Anthropic 32% 65%
DeepSeek 17% ~1%
OpenAI ~13% ~13%
Google (largest volume earlier in spring) n/p n/p

June 2026 (Vercel July 2026 index):

Provider Token share Spend share
Anthropic 32% 61%
Google ~24% n/p
DeepSeek 22.6% <4%
OpenAI 10.3% 16.1%

Two structural findings:

  1. Volume and value have decoupled. Anthropic takes roughly a third of tokens and ~61–65% of spend. DeepSeek takes a fifth to a quarter of tokens for under 4% of spend. Open-weight models as a group ran 29% of June tokens on under 4% of spend, up from 11% in April.
  2. The coding-agent slice is where the gap is widest. In May, within the AI-coding-agent use case, DeepSeek drove 49% of token volume but only 4% of cost, while Anthropic drove 28% of tokens and 70% of cost. In June, Anthropic still captured ≥72% of spend in every high-stakes use case, including coding agents, back-office agents and app generation.

Google's position is instructive: it leads consumer-shaped volume (57% of personal-assistant tokens in June) but runs <2% of Vercel coding-agent tokens. OpenAI, by contrast, is the most evenly distributed across use cases; its June cost-per-token rose ~50% relative to the market as GPT-5.x-class work shifted to more expensive reasoning.


4. Community usage: what developers install

SDK downloads are the community's "build with" signal. PyPI figures through Aug 2026 (vester.si AI Impact):

Python package Monthly downloads (Aug 2026) Note
litellm ~644M Multi-vendor calling layer, not a model provider
openai ~427M Also used as the interface for OpenAI-compatible providers
google-genai ~261M Rebranded Gemini SDK
anthropic ~197M Fastest direct-SDK growth: ~10x since Jan 2026
openai-agents ~35.8M OpenAI agent SDK
claude-agent-sdk ~31.6M Claude agent SDK; ~26x since Jan 2026

Ranking by raw installs: OpenAI first, Google second, Anthropic third — but three caveats matter for the coding-agent question:

  • The openai package is the de facto API-compatible interface for many providers (Groq, Together, Azure, etc.), so it overstates calls to OpenAI's own API. In May 2026, Presenc estimated OpenAI's combined Python+npm SDK total at ~380M/month vs Google ~275M and Anthropic ~191M (Presenc AI SDK Download Rankings).
  • On npm, the gap between OpenAI and Anthropic is much smaller: May 2026 showed openai at 84.4M vs @anthropic-ai/sdk at 71.6M monthly downloads — a 1.2x gap vs ~2.5x on PyPI.
  • The fastest-growing downloads in 2026 are not single-vendor SDKs at all. LiteLLM (~644M/month) and the agent SDKs (openai-agents, claude-agent-sdk) are the calling layer; agent applications increasingly route across providers rather than committing to one API key.

If you interpret "community usage" literally as package installs, OpenAI is the most popular provider. If you interpret it as who agent/coding software actually calls at runtime, downloads undercount the routing layer and overcount the interface default.


5. Enterprise and paid usage

Menlo Ventures (US enterprise, survey fielded Nov 2025)

The most recent large US enterprise survey (ZDNET/Menlo coverage) found:

  • Enterprise GenAI spend ~\(37B** in 2025; coding tools were the largest application category at **~\)4B.
  • Enterprise LLM spend share: Anthropic 40%, OpenAI 27%, Google 21%.
  • Coding-model share: Anthropic 54%, OpenAI 21%.
  • Only 16% of enterprise production systems qualified as true agents; the mainstream is still copilots and coding tools.

These are estimates from a VC survey (Menlo is an Anthropic investor) and are now ~9 months old, but they are the last broad US enterprise benchmark that separately measured coding.

Ramp AI Index (actual business payments, July 2026)

Payment-card data from 70,000+ US businesses (Crypto Briefing/Ramp coverage) tells a more current and more volatile story:

Metric (July 2026) Anthropic OpenAI
Share of tracked businesses paying for subscriptions or tokens 43.5% 39.7%
QoQ enterprise growth index 76 82
Flagship model share of that vendor's tokens Fable 5: 6% GPT-5.6 Sol: 25%
Flagship model share of that vendor's spend Fable 5: 11.4% GPT-5.6 Sol: 23%

Read carefully: Anthropic still leads in breadth (more companies buy it), while OpenAI is winning intensity (GPT-5.6 Sol users consume and spend more per customer) and growing faster in Q3. Ramp skews toward startups and mid-market businesses and measures who pays, not how much value they get.

Coding products themselves

Product-level tracking points the same way as the gateway data. TickerTrends' Aug 10 snapshot (covered in The LLM ARR Debate):

Product Tracked ARR (Aug 10, 2026) MoM growth
Claude Code ~$15.12B +5.2%
OpenAI Codex ~$8.83B +20.8%

Claude Code is roughly 1.7x Codex by tracked ARR, but Codex is growing four times faster month-over-month. On raw public-router volume, open-weight agent harnesses dominate, but the paid-codings product crown still belongs to Anthropic.


Pick the metric, get the answer:

Question #1 #2 Why
Most paid coding/agent inference (spend) Anthropic OpenAI 61–65% of Vercel gateway spend; ≥72% of high-stakes spend; 54% US enterprise coding share; Claude Code largest paid coding product
Most raw tokens through neutral routers DeepSeek OpenAI / Z.ai 19.0% of tracked OpenRouter tokens (Sep 4); ~half of Vercel coding-agent tokens on ~4% of cost
Most developer installs / API ecosystem OpenAI Google / Anthropic 427M monthly PyPI downloads of openai; 4M developers; largest direct API disclosed scale
Fastest-moving right now OpenAI (intensity), Z.ai/GLM & DeepSeek (volume) Codex ARR +20.8% MoM; GPT-5.6 Sol = 25% of OpenAI business tokens; GLM-5.3-Flash hit #1 model on OpenRouter within days

For the specific use case in the question — BYO API keys, LLM agents, coding inference — Anthropic is the most popular premium provider and the default where quality gates matter; DeepSeek is the most popular volume provider; OpenAI is the most popular ecosystem and the fastest-rising paid rival. Most serious agent stacks in 2026 do not choose one: they route routine coding to a cheap open-weight model and reserve Claude (or GPT-5.6-class) for the steps where a wrong answer is expensive.


Caveats

  • OpenRouter rankings count tokens, not requests, users, quality or spend, and they exclude private usage and first-party apps. Cross-provider token counts also reflect different tokenizers.
  • Vercel AI Gateway sees only traffic routed through its gateway and normalizes spend to list price; real negotiated prices differ.
  • SDK downloads include CI/bot/mirror noise and OpenAI-compatible interfaces.
  • Menlo and Ramp measure different populations; neither is a complete census of global inference.
  • The market moved materially within the past month: GLM-5.3-Flash (formerly "Ox Alpha") topped OpenRouter's weekly model list in late August, and Anthropic fell to <3% of OpenRouter top-20 model tokens by Aug 30. Rankings dated even two weeks ago may already be stale.

Sources