LLM Landscape — Latest Snapshot
A rolling digest, current as of early August 2026. Prices and benchmark values change quickly; the linked primary sources and dated snapshot pages are the authority for current values. Full data tables live in the snapshot archive below — this page is the quick orientation.
1. Landscape at a glance
Six releases dominated summer 2026. See the frontier model benchmark review for the full literature review and benchmark explainer.
| Tier | Models | One-line take |
|---|---|---|
| Frontier (closed) | Claude Fable 5, Claude Opus 5, GPT-5.6 Sol | Best aggregate benchmarks (Fable 5), best agentic coding (GPT-5.6 Sol), best per-dollar at premium tier (Opus 5) |
| Workhorse (closed) | GPT-5.6 Terra/Luna, Claude Sonnet 5 | Price-performance: Luna got an 80% price cut on Jul 30 |
| Frontier (open weights) | Kimi K3, GLM-5.2, DeepSeek V4 Flash 0731 | Kimi K3 leads open long-horizon agentic work; GLM-5.2 is the strongest open coding model; V4 Flash 0731 is the agentic cost leader |
Intelligence Index (Artificial Analysis, max-effort / adaptive reasoning):
| Rank | Model | Index |
|---|---|---|
| 1 | Kimi K3 (max) | 57 |
| 2 | Claude Opus 4.8 | 56 |
| 3 | GPT-5.6 Terra (max) | 55 |
| 4 | Claude Sonnet 5 | 53 |
| 5 | GPT-5.6 Luna (max) | 51 |
| 7 | GLM-5.2 (max) | 51 |
| 8 | DeepSeek V4 Flash 0731 | 50 |
| 10 | DeepSeek V4 Pro | 44 |
Note: Claude Fable 5 / GPT-5.6 Sol / Claude Opus 5 score higher (AA Intelligence Index ~58–60) but are evaluated on newer index versions; treat this table as the current-generation ranking where available.
Quick picks: - Coding / agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open) - Novel reasoning (ARC-AGI-3): Claude Opus 5 - Knowledge work: Claude Fable 5, GPT-5.6 Sol - Long context + multimodal: Claude Fable 5, Kimi K3, GPT-5.6 Sol - Price-performance: Claude Opus 5 (premium), GPT-5.6 Luna/Terra, DeepSeek V4 Flash 0731 (open)
2. Pricing at a glance
All prices USD per 1M tokens, cache-miss input / output, list price.
Frontier closed models
| Model | Input | Output | Context |
|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | 1M |
| Claude Opus 5 | $5.00 | $25.00 | 1M |
| GPT-5.6 Sol | $5.00 | $30.00 | 256K |
| GPT-5.6 Terra | $2.50 | $15.00 | 256K |
| GPT-5.6 Luna | $0.20 | $1.20 | 256K (after Jul 30 cut) |
| Claude Sonnet 5 | \(2.00–\)3.00 | \(10.00–\)15.00 | 1M |
Open-weight / cheap
| Model | Input | Output | Context |
|---|---|---|---|
| DeepSeek V4 Flash 0731 | $0.14 | $0.28 | 1M (cache-hit input $0.0028) |
| DeepSeek V4 Pro | $0.435 | $0.87 | 1M |
| GLM-5.2 | ~$1.12 | ~$3.92 | 1M |
| Kimi K3 | ~$2.80 | ~$14.00 | 1M |
| Muse Spark 1.1 | $1.25 | $4.25 | — |
Reference task cost (20,000 in + 5,000 out tokens, no cache)
| Model | Task cost |
|---|---|
| DeepSeek V4 Flash 0731 | $0.0042 |
| GPT-5.6 Luna | $0.0500 |
| Gemini 3.6 Flash | $0.0675 |
| Claude Sonnet 5 | $0.1350 |
| Claude Opus 4.8 | $0.2250 |
These are arithmetic estimates for a fixed token mix, not observed task costs. See Model intelligence & cost per task and the DeepSeek V4 Flash 0731 snapshot for the methodology and caveats.
3. Snapshot archive
Dated research snapshots with the full tables behind the digest above:
| Snapshot | Date | What it contains |
|---|---|---|
| LLM Landscape Snapshot (full) | 2026-07-19 | Full model list, Text/Agent Arena leaderboards, per-provider API pricing tables, value analysis |
| LLM Landscape Snapshot (week update) | 2026-07-26 | New releases (Claude Opus 5, Gemini Flash Cyber), product launches, incidents, funding |
| Chinese LLM API Pricing | 2026-07-26 | Token pricing for DeepSeek, Kimi K3, GLM-5.2, Grok 4.5, GPT-5.6, Gemini 3.6 Flash, Muse Spark 1.1 |
| DeepSeek V4 Flash 0731 — ranking & cost | 2026-08-01 | DeepSeek's efficiency-tier release: Intelligence Index, agent benchmarks, cost-per-task analysis |
| Model intelligence & cost per task | 2026-07-26 | Benchmark score vs token use vs task cost methodology and comparison |