Data sourced live via websearch on July 19, 2026. All prices in USD per million tokens unless noted. Benchmark dates reflect source snapshot dates.
1. Models in the Arena (as of July 2026)
1.1 Closed-Source / Frontier Models
Model
Provider
Latest Version
Context Window
Key Highlights
GPT-5.6 Sol
OpenAI
May 2026
256K
Flagship, 3 tiers: Sol/Terra/Luna
GPT-5.6 Terra
OpenAI
May 2026
256K
Mid-tier 5.6
GPT-5.6 Luna
OpenAI
May 2026
256K
Budget 5.6
GPT-5.5
OpenAI
Early 2026
256K
Previous-gen flagship
GPT-5.5 Pro
OpenAI
Early 2026
256K
Deep research / highest reasoning
GPT-5.4
OpenAI
2025
256K
Two gens back, still available
GPT-5.4 Mini
OpenAI
2025
128K
Small/fast
GPT-5.4 Nano
OpenAI
2025
128K
Ultra-cheap
Claude Fable 5
Anthropic
~Jun 2026
1M
Text+Agent Arena #1, premium tier
Claude Mythos 5
Anthropic
~Jun 2026
1M
Limited availability, same pricing as Fable 5
Claude Opus 4.8
Anthropic
~Jun 2026
1M
Current Opus, fast mode available
Claude Opus 4.7
Anthropic
Early 2026
1M
Previous Opus
Claude Sonnet 5
Anthropic
~Jun 2026
1M
Current Sonnet, intro pricing through Aug 31
Claude Sonnet 4.6
Anthropic
Early 2026
1M
Previous Sonnet
Claude Haiku 4.5
Anthropic
Early 2026
1M
Budget lightweight
Gemini 3 Pro
Google
2026
1M+
Top Google model, strong vision
Gemini 3.5 Flash
Google
2026
1M
Mid-tier Flash
Gemini 3.1 Pro Preview
Google
2026
1M
Preview model
Grok 4.5
xAI
~Jun 2026
500K
Latest Grok, xAI flagship
Grok 4.3
xAI
Early 2026
1M
Previous-gen Grok
DeepSeek V4 Pro (cloud API)
DeepSeek
2026
1M
Latest DeepSeek, cloud API
DeepSeek V4 Flash (cloud API)
DeepSeek
2026
1M
Budget DeepSeek, ultra-cheap
1.2 Open-Weight Models
Model
Provider
Context Window
License
Key Highlights
DeepSeek V4 Pro
DeepSeek
1M
MIT
Strong reasoning, thinking mode
DeepSeek V4 Flash
DeepSeek
1M
MIT
Ultra-cheap, fast
Qwen 3.7 Max
Alibaba (Qwen Team)
128K
Apache 2.0
Latest Qwen generation
Qwen 3.7 Plus
Alibaba (Qwen Team)
128K
Apache 2.0
Budget Qwen
Gemma 4 31B
Google
128K
Apache 2.0
Google open model
GLM 5.2 (Max)
Z.ai
128K
MIT
Top Chinese open model, agent-strong
GLM 5.1
Z.ai
128K
MIT
Previous GLM
Kimi K3
Moonshot
128K
Modified MIT
WebDev #1
Kimi K2.7 Code
Moonshot
128K
Modified MIT
Code-specialized
Mimo V2.5 Pro
Xiaomi
128K
MIT
Emerging challenger
Minimax M3
MiniMax
128K
Community
Growing presence
Nemotron 3 Ultra
Nvidia
128K
OpenMDW-1.1
Nvidia open-weight entry
1.3 Reasoning / Thinking Models
Model
Provider
Base Model
Key Differentiator
Claude Opus 4.8 Thinking
Anthropic
Opus 4.8
Agent Arena #3
Claude Opus 4.7 Thinking
Anthropic
Opus 4.7
Text Arena #3
Claude Opus 4.6 Thinking
Anthropic
Opus 4.6
Text Arena #2
DeepSeek V4 Pro Thinking
DeepSeek
V4 Pro
Strong math/logic
DeepSeek V4 Flash Thinking
DeepSeek
V4 Flash
Budget reasoning
GPT-5.6 Sol xHigh
OpenAI
GPT-5.6 Sol
Agent Arena #2, reasoning-effort tiers
GPT-5.5 xHigh
OpenAI
GPT-5.5
Agent Arena #4
Grok 4.3 Think
xAI
Grok 4.3
Budget reasoning
1.4 Notable Small / Efficient Models
Model
Provider
Use Case
GPT-5.4 Nano
OpenAI
Ultra-cheap chat, \(0.20/\)1.25 per MTok
Claude Haiku 4.5
Anthropic
Budget agent, \(1/\)5 per MTok
Gemma 4 26B
Google
On-device capable
DeepSeek V4 Flash (cache hit)
DeepSeek
$0.0028/MTok, extreme cost efficiency
Grok Build 0.1
xAI
Code/developer specialized, \(1/\)2
2. Leaderboards & Benchmarks
2.1 LMSys Text Arena (Overall Elo)
Snapshot date: July 19, 2026. Source: https://lmarena.ai
Rank
Model
Overall Elo
Expert
Hard Prompts
Coding
Math
Creative Writing
Instruction Following
Longer Query
1
Claude Fable 5
1507 ±7
1
1
1
1
1
1
1
2
Claude Opus 4.6 Thinking
1504 ±4
2
2
3
2
2
2
2
3
Claude Opus 4.7 Thinking
1503 ±4
5
4
1
5
3
3
4
4
Claude Opus 4.6
1498 ±4
4
3
4
6
7
4
3
5
Claude Opus 4.7
1494 ±4
3
5
5
10
5
6
5
6
Muse Spark 1.1 (Meta)
1493 ±8
24
7
8
14
18
12
21
7
Muse Spark (Meta)
1487 ±6
37
12
13
38
13
26
36
8
Gemini 3 Pro (Google)
1486 ±4
22
13
23
20
4
17
13
9
Kimi K3
1486 ±11
18
10
10
-
6
9
8
10
GPT-5.6 Sol xHigh
1486 ±9
8
11
14
24
8
10
15
11
Gemini 3.1 Pro Preview
1485
12
9
18
12
9
8
10
12
Claude Opus 4.8 Thinking
1483
6
6
6
7
14
5
6
13
GPT-5.5 High
1481
9
15
21
13
26
14
18
14
GPT-5.4 High
1478
7
17
19
9
36
15
23
15
Gemini 3.5 Flash High
1474
16
26
43
3
12
28
27
2.2 Agent Arena (LMSys Agent Leaderboard)
Snapshot date: July 12, 2026. Source: https://lmarena.ai/leaderboard/agent
Rank
Model
Net Improvement
Confirmed Success
Praise vs Complaint
Steerability
Bash Recovery
Tool Hallucination
Sessions
1
Claude Fable 5 (High)
13.94% ±1.56%
17.27% ±2.75%
30.65% ±5.67%
12.07% ±2.94%
8.39% ±1.81%
1.33% ±0.15%
16,059
2
GPT-5.6 Sol (xHigh)
10.94% ±3.76%
10.93% ±3.73%
17.64% ±7.48%
17.26% ±16.35%
7.53% ±1.68%
1.33% ±0.15%
7,881
3
Claude Opus 4.8 (Thinking)
9.28% ±1.35%
7.82% ±2.63%
17.96% ±4.91%
10.15% ±2.48%
9.91% ±1.00%
0.56% ±0.71%
33,392
4
GPT-5.5 (xHigh)
8.26% ±0.87%
6.88% ±1.71%
12.61% ±3.14%
6.70% ±1.70%
13.79% ±0.77%
1.33% ±0.15%
36,289
5
Claude Sonnet 5 (High)
8.00% ±1.71%
10.33% ±3.39%
13.90% ±6.50%
4.92% ±3.24%
9.62% ±0.88%
1.21% ±0.16%
23,640
6
Claude Opus 4.7 (Thinking)
7.73% ±1.22%
5.22% ±2.56%
11.36% ±4.35%
9.17% ±2.31%
11.70% ±1.16%
1.22% ±0.17%
34,357
7
Claude Opus 4.7
7.63% ±1.22%
5.28% ±2.54%
13.38% ±4.39%
7.56% ±2.27%
10.66% ±1.32%
1.27% ±0.16%
34,950
8
GPT-5.5 (High)
7.16% ±0.75%
6.42% ±1.47%
7.61% ±2.61%
8.72% ±1.42%
11.75% ±1.00%
1.33% ±0.15%
61,455
9
GLM 5.2 (Max)
6.24% ±1.10%
8.72% ±2.15%
12.59% ±4.01%
4.10% ±2.03%
4.45% ±1.06%
1.33% ±0.15%
31,993
10
GPT-5.5
6.05% ±0.72%
4.33% ±1.46%
5.53% ±2.51%
7.99% ±1.33%
11.06% ±0.81%
1.33% ±0.15%
62,335
19
Gemini 3.1 Pro Preview
-0.70% ±0.69%
-0.63% ±1.53%
0.07% ±2.19%
1.85% ±1.26%
7.21% ±1.15%
1.29% ±0.15%
61,459
22
DeepSeek V4 Pro
-1.11% ±1.36%
-2.95% ±3.30%
-4.79% ±4.46%
3.10% ±2.94%
4.50% ±1.19%
0.80% ±0.34%
10,335
31
Gemini 3 Flash
-8.54% ±0.77%
-8.46% ±1.58%
-12.83% ±1.90%
5.24% ±1.27%
16.46% ±2.09%
0.28% ±1.26%
62,100
35
Grok 4.3
-15.31% ±1.01%
-10.12% ±1.62%
-16.64% ±1.86%
8.96% ±1.24%
41.90% ±4.11%
1.10% ±0.17%
61,497
2.3 SWE-bench Verified
Data not available via static HTML scrape (requires JavaScript). Source: https://www.swebench.com/verified.html
Rank
Model
Resolved Rate (%)
Date
[Pending — JS-required leaderboard. Refer to swebench.com for live data.]
2.4 AIME 2025 (Math Reasoning)
Model
AIME Score
Date
[Pending — dedicated search required. See Text Arena Math column for relative ranking.]
2.5 GPQA Diamond (Graduate-Level Q&A)
Model
GPQA Diamond Score
Date
[Pending — dedicated search required.]
2.6 MMLU-Pro
Model
MMLU-Pro Score
Date
[Pending — dedicated search required.]
2.7 HumanEval+ / BigCodeBench
Model
HumanEval+ Score
BigCodeBench Score
Date
[Pending — dedicated search required. See Text Arena Coding column for relative ranking.]
3. API Pricing Snapshot
All prices per million tokens (USD). Sourced from official provider pricing pages on July 19, 2026.
3.1 OpenAI
Model
Input
Cached Input
Output
Batch Input
Batch Output
Max Context
GPT-5.6 Sol (short ctx)
$5.00
$0.50
$30.00
$2.50
$15.00
256K
GPT-5.6 Sol (long ctx)
$10.00
$1.00
$45.00
$5.00
$22.50
256K
GPT-5.6 Terra (short ctx)
$2.50
$0.25
$15.00
$1.25
$7.50
256K
GPT-5.6 Terra (long ctx)
$5.00
$0.50
$22.50
$2.50
$11.25
256K
GPT-5.6 Luna (short ctx)
$1.00
$0.10
$6.00
$0.50
$3.00
256K
GPT-5.6 Luna (long ctx)
$2.00
$0.20
$9.00
$1.00
$4.50
256K
GPT-5.5 (short ctx)
$5.00
$0.50
$30.00
$2.50
$15.00
256K
GPT-5.5 (long ctx)
$10.00
$1.00
$45.00
$5.00
$22.50
256K
GPT-5.5 Pro
$30.00
-
$180.00
$15.00
$90.00
256K
GPT-5.4
$2.50
$0.25
$15.00
$1.25
$7.50
256K
GPT-5.4 Mini
$0.75
$0.075
$4.50
$0.375
$2.25
128K
GPT-5.4 Nano
$0.20
$0.02
$1.25
$0.10
$0.625
128K
GPT-5.4 Pro
$30.00
-
$180.00
$15.00
$90.00
256K
Priority processing: 2x standard pricing. Flex processing: same as batch pricing. Data residency: +10% uplift for models released on/after March 5, 2026.
3.2 Anthropic
Model
Input
Output
5m Cache Write
1h Cache Write
Cache Read
Batch Input
Batch Output
Claude Fable 5
$10.00
$50.00
$12.50
$20.00
$1.00
$5.00
$25.00
Claude Mythos 5
$10.00
$50.00
$12.50
$20.00
$1.00
$5.00
$25.00
Claude Opus 4.8
$5.00
$25.00
$6.25
$10.00
$0.50
$2.50
$12.50
Claude Opus 4.7
$5.00
$25.00
$6.25
$10.00
$0.50
$2.50
$12.50
Claude Opus 4.6
$5.00
$25.00
$6.25
$10.00
$0.50
$2.50
$12.50
Claude Sonnet 5 (intro, through Aug 31)
$2.00
$10.00
$2.50
$4.00
$0.20
$1.00
$5.00
Claude Sonnet 5 (from Sep 1)
$3.00
$15.00
$3.75
$6.00
$0.30
$1.50
$7.50
Claude Sonnet 4.6
$3.00
$15.00
$3.75
$6.00
$0.30
$1.50
$7.50
Claude Haiku 4.5
$1.00
$5.00
$1.25
$2.00
$0.10
$0.50
$2.50
Note: Models 4.8+ use a new tokenizer producing ~30% more tokens for the same text. Batch API: 50% discount. Fast mode pricing: Claude Opus 4.8 \(10/\)50, Claude Opus 4.7 \(30/\)150. All Opus 4.8/Sonnet 5 models support the full 1M context at standard pricing.Data residency: +10% uplift for US-only inference (Opus 4.6+ and Sonnet 4.6+).
3.3 xAI (Grok)
Model
Input (< 200K)
Cached Input
Output (< 200K)
Input (≥ 200K)
Output (≥ 200K)
Max Context
Grok 4.5
$2.00
$0.30
$6.00
$4.00
$12.00
500K
Grok 4.3
$1.25
$0.20
$2.50
$2.50
$5.00
1M
Grok 4.20 Reasoning
$1.25
$0.20
$2.50
$2.50
$5.00
1M
Grok 4.20 Non-Reasoning
$1.25
$0.20
$2.50
$2.50
$5.00
1M
Grok Build 0.1
$1.00
$0.20
$2.00
$2.00
$4.00
256K
3.4 DeepSeek
Model
Input (Cache Miss)
Input (Cache Hit)
Output
Context Window
Max Output
DeepSeek V4 Flash
$0.14
$0.0028
$0.28
1M
384K
DeepSeek V4 Pro
$0.435
$0.003625
$0.87
1M
384K
Both support thinking and non-thinking modes. Concurrency: V4 Flash 2500, V4 Pro 500.
3.5 Google Gemini
Model
Input
Output
Context Window
[Pending — Google AI pricing page timed out. Refer to ai.google.dev/pricing for live data.]
4. Cost-per-Intelligence Ratio (Quick Heuristic)
Normalized score = Text Arena Elo normalized to 0-100 scale (baseline 1300-1550). Cost estimate: standard task of 500 input + 2000 output tokens.
Model
Norm. Score
Cost per Task (500+2000 tokens)
Value Ratio (Score / Cost × 1000)
Claude Fable 5
100.0
$0.105
952
Claude Opus 4.8
95.6
$0.0525
1821
GPT-5.6 Sol
93.2
$0.0625
1491
Claude Sonnet 5
82.0
$0.021
3905
GPT-5.6 Terra
85.0
$0.03125
2720
Claude Opus 4.7
93.6
$0.0525
1783
Muse Spark 1.1
93.2
N/A (no public API pricing)
-
GPT-5.6 Luna
78.0
$0.013
6000
Grok 4.5
82.0
$0.013
6308
DeepSeek V4 Pro
68.0
$0.00196
34700
DeepSeek V4 Flash
64.0
$0.00063
101587
GPT-5.4 Nano
60.0
$0.0026
23077
Claude Haiku 4.5
72.0
$0.0105
6857
Note: DeepSeek pricing assumes cache miss. Cache hit pricing would make value ratios even more extreme. Google, Grok 4.5 pricing does not fully account for long-context surcharges.
5. Summary Output
5.1 Top 10 Models by Text Arena Elo
Rank
Model
Provider
Elo
Release Date
1
Claude Fable 5
Anthropic
1507 ±7
~Jun 2026
2
Claude Opus 4.6 Thinking
Anthropic
1504 ±4
Early 2026
3
Claude Opus 4.7 Thinking
Anthropic
1503 ±4
Early 2026
4
Claude Opus 4.6
Anthropic
1498 ±4
Early 2026
5
Claude Opus 4.7
Anthropic
1494 ±4
Early 2026
6
Muse Spark 1.1
Meta
1493 ±8
~Jun 2026
7
Muse Spark
Meta
1487 ±6
2026
8
Gemini 3 Pro
Google
1486 ±4
2026
9
Kimi K3
Moonshot
1486 ±11
2026
10
GPT-5.6 Sol xHigh
OpenAI
1486 ±9
May 2026
5.2 Top 5 Models by Agent Arena
Rank
Model
Net Improvement
Date
1
Claude Fable 5 (High)
13.94%
Jul 2026
2
GPT-5.6 Sol (xHigh)
10.94%
Jul 2026
3
Claude Opus 4.8 (Thinking)
9.28%
Jul 2026
4
GPT-5.5 (xHigh)
8.26%
Jul 2026
5
Claude Sonnet 5 (High)
8.00%
Jul 2026
5.3 New Entrants Since Early 2026
Model
Provider
Release Date
Significance
Claude Fable 5
Anthropic
~Jun 2026
New tier above Opus, #1 in both main Arenas
Claude Mythos 5
Anthropic
~Jun 2026
Limited availability, same-level as Fable
GPT-5.6 series (Sol/Terra/Luna)
OpenAI
May 2026
3-tier pricing, closing the gap with Anthropic
Muse Spark / Spark 1.1
Meta
~Jun 2026
Meta's strongest model yet, #6-7 Text Arena
Claude Sonnet 5
Anthropic
~Jun 2026
Successor to Sonnet 4.6, intro pricing
Claude Opus 4.8
Anthropic
~Jun 2026
Fast mode, refined Opus
Grok 4.5
xAI
~Jun 2026
Significant upgrade from Grok 4.3
DeepSeek V4 (Flash/Pro)
DeepSeek
2026
New generation, retains extreme affordability
GLM 5.2 (Max)
Z.ai
2026
Chinese open model competitive with frontier
5.4 Biggest Surprises / Trends
Trend
Description
Evidence
Anthropic's two-tier premium
Fable 5 (\(10/\)50) is priced 2× above Opus 4.8 (\(5/\)25), creating a new super-premium category
Pricing pages
OpenAI's tiered GPT-5.6 strategy
Three model levels (Sol/Terra/Luna) at \(5/\)2.50/$1 input — same base model, different compute budgets
API pricing
Meta seriously competitive
Muse Spark 1.1 at #6 Text Arena is Meta's best showing ever, blurring open/closed gap
Arena rankings
Agent gap is still wide for non-frontier
Top 10 Agent Arena all from Anthropic or OpenAI; Google Gemini 3.1 ranks 19th, DeepSeek V4 Pro at 22nd
Agent leaderboard
An entire tier of capable Chinese models
GLM 5.2, Qwen 3.7, Kimi K3, Mimo V2.5, Minimax M3 all rank in top 40 and are available with permissive licenses
Arena rankings
DeepSeek caching is absurdly cheap
$0.0028/MTok input on cache hit makes V4 Flash the cheapest frontier-capable model by 50×+
DeepSeek API docs
Claude dominates categories 7/7
Fable 5 takes the #1 spot in every Text Arena category — Expert, Hard Prompts, Coding, Math, Creative Writing, Instruction Following, Longer Query
LMSys Text Arena
5.5 Best Value Models (Intelligence per Dollar)
Model
Rationale
DeepSeek V4 Flash
Frontier-capable, \(0.14/\)0.28 per MTok, $0.0028 on cache hit. 50-100× cheaper than Claude Fable 5 for many tasks
Claude Sonnet 5 (intro pricing)
\(2/\)10, Agent Arena top-5, intro price through Aug 31
GPT-5.6 Luna
\(1/\)6, reasonable intelligence from the GPT-5.6 family
All prices confirmed from official sources on July 19, 2026. Verify before citing.
Arena Elo scores are dynamic; this snapshot captures one point in time.
SWE-bench Verified, AIME 2025, GPQA Diamond, and MMLU-Pro scores require dedicated searches and may not be available on the same day.
Anthropic's new tokenizer (Opus 4.8+, Sonnet 5+) generates ~30% more tokens for the same text — this should be factored into any cost-per-task calculation.
DeepSeek V4 Flash cache hit pricing ($0.0028/MTok) makes it the overwhelming value leader for repeated-context workloads.
Google Gemini pricing was unavailable at snapshot time (timeout). Refer to https://ai.google.dev/pricing for current rates.