Skip to content

Frontier LLM Releases of Summer 2026: A Benchmark-Focused Review

A literature review of six official model releases — Claude Fable 5, Claude Opus 5, GPT-5.6, DeepSeek-V4-Flash-0731, Kimi K3, and GLM-5.2 — with a benchmark explainer and an executive summary.


Part 1 — Literature Review: What the Official Releases Claim

All findings below are taken from the official release announcements (URLs in the references section). A recurring theme in 2026: benchmarks have shifted from static knowledge tests (MMLU, GSM8K) to agentic, long-horizon evaluations — coding agents, terminal use, knowledge work, and computer use — with cost-per-task as a first-class metric alongside accuracy.

1.1 Claude Fable 5 (Anthropic, June 9, 2026; redeployed July 1)

Fable 5 is a "Mythos-class" model made safe for general use — the same underlying model as Claude Mythos 5, with heavy safety classifiers (cyber/bio/chemistry requests fall back to Opus 4.8; <5% of sessions affected). Anthropic claims it is "state-of-the-art on nearly all tested benchmarks," with the largest lead on long, complex tasks.

Benchmarks cited: FrontierCode (Cognition) — highest score among frontier models; FrontierBench — highest score; CursorBench — state of the art (per Cursor); Hebbia Finance Benchmark — highest score; ViBench (Replit) — highest. Cross-cited third-party results: GDPval-AA 1,759.6 Elo (best), AA Intelligence Index 59.9 (best), AA Coding Agent Index 77.2, SWE-bench Pro 80%, Terminal-Bench 2.1 83.1, BrowseComp 84.3%, Toolathlon 61.7, GPQA Diamond 92.6%, FrontierMath Tier 4 87.8% (best), Agents' Last Exam 40.5%, HealthBench Professional 60.9%, GraphWalks BFS 1M 79.4.

Context: Suspended June 12–30 due to US export controls following an Amazon jailbreak report; redeployed July 1 globally. Pricing \(10/\)50 per M tokens. Early testimonials emphasize months-of-engineering-in-days workloads (Stripe's 50M-line migration) and best-in-class vision (Pokémon FireRed with vision-only harness).

1.2 Claude Opus 5 (Anthropic, July 24, 2026)

Positioned as "near-Fable intelligence at half the price" and the new default on Claude Max. Anthropic claims new state-of-the-art on coding and knowledge-work evaluations.

Benchmarks cited: Frontier-Bench v0.1 — surpasses all models, more than doubles Opus 4.8's performance at lower cost-per-task; CursorBench 3.2 — within 0.5% of Fable 5 at half cost; ARC-AGI 3 — 3× the next-best model; Zapier AutomationBench — ~1.5× next-best pass rate at equal cost; OSWorld 2.0 — outperforms every model at any given cost; GDPval-AA v2 — new SOTA; HLE — best and most cost-efficient; also AutomationBench, DeepSearchQA, AA Coding Agent Index, FrontierCode 1.1 (Devin's eval), OSS-Fuzz (cyber: close to Mythos 5 at finding, far behind at exploiting), and internal life-sciences evals (spectroscopy +10.2pp over Opus 4.8, protein variant prediction +7.7pp). Priced \(5/\)25 per M tokens (same as Opus 4.8).

1.3 GPT-5.6 Sol / Terra / Luna (OpenAI, July 9, 2026 GA; July 30 price cut)

A three-tier family: Sol (flagship), Terra (balanced), Luna (cheapest). OpenAI's framing is "performance per dollar": Sol is claimed to be "state-of-the-art across coding, knowledge work, cybersecurity, and science while using fewer tokens at lower estimated cost." New ultra effort setting coordinates 4 parallel agents.

Benchmarks cited (from the release table): Agents' Last Exam 52.7% (53.6 claimed) vs Fable 5's 40.5%; AA Coding Agent Index v1.1 80 (best); Terminal-Bench 2.1 88.8 (91.9 with ultra); DeepSWE v1.1 72.7; SWE-bench Pro 64.6; BrowseComp 90.4 (92.2 ultra); OSWorld 2.0 62.6; GPQA Diamond 94.6; FrontierMath 89% (Tiers 1–3), 83% (Tier 4); MMMU Pro 84.6%; gdp.pdf 30.7%; GDPval-AA v2 1,747.8 Elo; AA Intelligence Index 58.9; HealthBench Professional 60.5%; GeneBench Pro 28.7%; LifeSciBench 59.9%; AutomationBench 18.1%; Toolathlon 58%; OpenAI MRCR v2 91.5% (256K–512K); GraphWalks BFS 79.4% (1M); ARC-AGI-3 7.78% (vs Opus 4.8's 1.5%); cybersecurity — ExploitBench 73.5% (vs GPT-5.5's 47.9%), ExploitGym 33.7%, SEC-Bench Pro 71.2%, CTF 96.7%; self-improvement — RSI Index 57.9, KernelGen 1P 61.1, NanoGPT 9.69%. Terra/Luna beat Fable 5 on Agents' Last Exam at ~1/16th the cost. July 30: Luna price cut 80%, Terra 20% (Luna now \(0.20/\)1.20 per M tokens).

1.4 DeepSeek-V4-Flash-0731 (DeepSeek, July 31, 2026)

Official release of the V4-Flash API (public beta), a re-post-trained version of V4-Flash-Preview with unchanged architecture/size. V4-Pro official release is pending. DeepSeek's headline: "significantly enhanced agent capabilities, benchmark results far exceeding V4-Pro-Preview."

Benchmarks cited: Terminal Bench 2.1 82.7; NL2Repo 54.2; Cybergym 76.7; DeepSWE 54.4; Toolathlon verified 70.3; Agent Last Exam 25.2; Automation Bench (Public) 25.1; DSBench-FullStack (internal) 68.7; DSBench-Hard (internal) 59.6. Evaluated with DeepSeek's own "Harness minimal mode" at max effort. Native Responses API support, adapted for Codex.

1.5 Kimi K3 (Moonshot AI, July 16–17, 2026)

The world's first open 2.8-trillion-parameter model (MoE, 16/896 active), native multimodal, 1M-token context, open weights released by July 27. Architecture: Kimi Delta Attention + Attention Residuals + Stable LatentMoE (~2.5× scaling efficiency vs K2). Moonshot's own verdict: "overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol," while outperforming all other tested models.

Benchmarks cited: DeepSWE v1.1 67.3 (official leaderboard, mini-SWE-agent harness); BrowseComp 90.4 (at 1M context, no context management); Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, MLS Bench Lite, KCB 2.0 (in-house), OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench (600-task public subset), GDPval-AA / AA-Briefcase / APEX-Agents (cited from Artificial Analysis), MMMU-Pro, PerceptionBench (Moonshot's own atomic-vision benchmark), ZeroBench. Open-ended capability demos: GPU kernel optimization (competitive with Fable 5), building a Triton-like compiler (MiniTriton) from scratch, designing a chip in a 48-hour autonomous run, and reproducing I–Love–Q astrophysics results (~2 hours vs 1–2 weeks human). Known limitations acknowledged: sensitivity to thinking-history, excessive proactiveness.

1.6 GLM-5.2 (Z.AI / Zhipu, June 13, 2026; open weights June 16)

744B-A40B MoE, 1M-token context, MIT license, built on the IndexShare sparse-attention architecture (2.9× FLOPs reduction at 1M context). Claims: "strongest open-source model on standard coding benchmarks" and a substantial leap in long-horizon capability over GLM-5.1.

Benchmarks cited (official blog; suite cross-confirmed by Kimi K3's footnotes): Terminal-Bench 2.1 81.0 (vs GLM-5.1's 63.5; within a few points of Claude Opus 4.8's 85.0, ahead of Gemini 3.1 Pro); SWE-bench Pro 62.1 (vs 58.4); plus DeepSWE, Program Bench, SWE Marathon, MLS Bench Lite, KCB 2.0, OfficeQA Pro, SpreadsheetBench 2, AutomationBench, BrowseComp, GDPval-AA. GLM-5.1 (its predecessor in the same family) held SOTA on SWE-bench Pro and NL2Repo at its release.


Part 2 — Benchmark Explainer: What Each Evaluation Tests

2.1 Software Engineering & Agentic Coding

Benchmark What it tests Who runs it
Terminal-Bench 2.1 Real command-line terminal usage: agents run shell commands, navigate file systems, use CLI tools to complete dev/ops tasks TAU-bench community / AA (tbench.org)
SWE-bench Pro Resolving real GitHub issues end-to-end in real codebases (harder than SWE-bench Verified; includes harder repos) OpenAI + SWE-bench team
DeepSWE v1.1 Long-horizon software engineering in large real-world codebases, sustaining work over many steps datacurve.ai
Frontier-Bench v0.1 Long-horizon engineering tasks incl. unusual tooling (e.g., rebuild a part as a FreeCAD 3D model), run on mini-SWE-agent harness; mean reward over 5 attempts Anthropic internal (per release)
FrontierCode 1.1 Hard coding tasks graded against production-quality codebase standards (code review, tests) Cognition
CursorBench 3.2 Agent performance inside the Cursor IDE workflow (edits, tests, iterations) Cursor
AA Coding Agent Index v1.1 Composite index of coding-agent performance: implementation, terminal use, real codebases Artificial Analysis (independent)
NL2Repo Generating a complete working repository from a natural-language spec NL2Repo community
Program Bench Program synthesis: generating correct code to spec across many languages Vals AI
SWE Marathon Marathon-style multi-hour engineering sessions in real repos swe-marathon.org
FrontierSWE Frontier-difficulty SWE tasks, dominance scoring frontierswe.com
PostTrain Bench Post-training/RL quality: how well models perform after reward-modeling phases posttrainbench.com
MLS Bench Lite Machine-learning systems engineering (training, tuning, infra) MLS Bench
KCB 2.0 Kimi's in-house coding benchmark (full-stack + agentic) Moonshot AI (internal)
DSBench-FullStack / -Hard DeepSeek's internal full-stack development and hard coding-agent test sets DeepSeek (internal)
KernelGen 1P / NanoGPT / RSI Index Self-improvement: kernel optimization, tiny training runs, aggregate recursive-self-improvement score OpenAI (internal)
Internal Research Debugging Eval Debugging research systems, optimizing training recipes OpenAI (internal)

2.2 Knowledge Work & Agentic Productivity

Benchmark What it tests Who runs it
Agents' Last Exam Long-running professional workflows across 55 fields, end-to-end agentic tasks agents-last-exam.org
GDPval-AA Predicting GDP per capita from economic documents — benchmark of analytical/reasoning strength Artificial Analysis
AA Intelligence Index v4.1 Broad composite: agentic work, coding, scientific reasoning, general capabilities Artificial Analysis
Big Finance Bench Financial analysis tasks (multi-hop research, quantitative reasoning) Model ML / Rogo
Management Consulting Tasks Internal consulting-style analysis workflows OpenAI (internal)
AutomationBench (Zapier) Completing business automation workflows end-to-end via tool chains Zapier
OfficeQA Pro Office/document tasks over PDF corpora rendered as images (no text layer) Community (per Kimi)
SpreadsheetBench 2 Spreadsheet formula/chart/analysis tasks Community (per Kimi)
BrowseComp Agentic web browsing: multi-step research with tool use OpenAI
OSWorld 2.0 Computer use: operating GUIs, clicking, typing, using real apps OSWorld project
MCP Atlas Tool use over MCP servers (500-task public subset, Gemini judge) MCP community
Toolathlon General tool-calling accuracy across many tools Toolathlon project
DeepSearchQA Deep multi-hop search + QA Anthropic (internal)
APEX-Agents / AA-Briefcase Agentic knowledge-work evals from Artificial Analysis Artificial Analysis
Management Consulting / Finance evals (Hebbia, IMC) Senior-level financial reasoning, document/chart interpretation Hebbia, IMC

2.3 Reasoning & Science

Benchmark What it tests Who runs it
GPQA Diamond Graduate-level, "Google-proof" multiple-choice questions in physics/chemistry/biology; experts score ~65%, expert non-specialists 34% idavidrein/gpqa
FrontierMath (v2) Original frontier mathematics problems, graded by difficulty tiers (Tier 1–3 / Tier 4) Epoch AI
HLE (Humanity's Last Exam) 2,500 expert-written frontier questions across 100+ subjects, multimodal; published in Nature CAIS + Scale AI
ARC-AGI-3 Novel abstraction/reasoning puzzles with core-knowledge priors — "easy for humans, hard for AI" ARC Prize
HealthBench Professional Professional clinical/medical reasoning Community (per OpenAI)
GeneBench Pro Long-horizon genomics & quantitative-biology analysis OpenAI
LifeSciBench Real-world biology and life-science research workflows Community (per OpenAI)
MedChemBench Medicinal-chemistry reasoning (internal) OpenAI (internal)
gdp.pdf Parsing/analyzing complex PDF documents (visual layout + content) Artificial Analysis
FrontierMath Tier 4 The hardest tier of original math research problems Epoch AI

2.4 Multimodal

Benchmark What it tests Who runs it
MMMU Pro College-level multimodal understanding (images + text) across art, science, engineering; with/without tool use MMMU team
PerceptionBench Atomic visual perception (Kimi's own benchmark for fine-grained vision) Moonshot AI
ZeroBench Zero-shot hard multimodal reasoning tasks Community

2.5 Long Context

Benchmark What it tests Who runs it
OpenAI MRCR v2 Multi-round context recall with 8 needles at 256K–512K and 512K–1M token windows OpenAI
GraphWalks BFS BFS traversal over graph structures at 256K and 1M tokens — deep long-context reasoning Community (per OpenAI)

2.6 Cybersecurity (frontier-tier models only)

Benchmark What it tests Who runs it
ExploitBench Progress from reaching vulnerable code to arbitrary code execution OpenAI
ExploitGym Turning real-world vulnerabilities into working exploits under time caps (2h/6h) OpenAI
SEC-Bench Pro Proof-of-concept generation on complex software OpenAI/community
Capture-the-Flag Solving CTF challenges (offensive security breadth) Community
OSS-Fuzz Finding then exploiting vulnerabilities in real open-source code Anthropic (built on Google's OSS-Fuzz)
CyberGym Reproducing target vulnerabilities (public leaderboard metric) CyberGym project
CyScenarioBench Success across realistic cyber scenarios Anthropic

Part 3 — Executive Summary: What Each Model Is Good At

Cross-model comparison on shared benchmarks (official numbers)

Benchmark Claude Fable 5 Claude Opus 5 GPT-5.6 Sol DeepSeek V4-Flash-0731 Kimi K3 GLM-5.2
Terminal-Bench 2.1 83.1 88.8 (91.9 ultra) 82.7 ~85 (Kimi harness) 81.0
DeepSWE v1.1 69.7 72.7 54.4 67.3 ✓ (per blog)
SWE-bench Pro 80.0 64.6 62.1
BrowseComp 84.3 90.4 (92.2 ultra) 90.4 ✓ (per blog)
Agents' Last Exam 40.5 52.7 25.2
GDPval-AA v2 (Elo) 1,759.6 new SOTA 1,747.8 ✓ (per blog)
AA Coding Agent Index 77.2 80
GPQA Diamond 92.6 94.6
ARC-AGI-3 3× next-best 7.78
Toolathlon 61.7 58.0 70.3 (verified)
OSWorld 2.0 54.8 (Opus 4.8) best per-cost 62.6

(— = not reported in the releases we reviewed; ✓ = reported in the release blog but exact figure was image-only or unextractable. Kimi K3 scores used its own harness unless noted.)

3.1 Claude Fable 5 — the all-round frontier benchmark

Best aggregate benchmark coverage of any model at release: GDPval-AA (1,759.6 Elo), AA Intelligence Index (59.9), FrontierMath Tier 4 (87.8), GraphWalks 1M long-context (79.4), HealthBench (60.9), top scores on FrontierCode, CursorBench, Hebbia Finance. Excels at vision (native multimodal, state-of-the-art per Anthropic) and very long autonomous tasks where its lead widens. Weaknesses: expensive (\(10/\)50), safety classifiers cause fallbacks, and it loses head-to-head on Agents' Last Exam (40.5 vs 52.7), ARC-AGI-3, and GPQA Diamond vs GPT-5.6 Sol.

3.2 Claude Opus 5 — the efficiency frontier at Opus tier

Not the absolute best at anything single-metric, but the best per dollar: Frontier-Bench v0.1 SOTA (2× Opus 4.8 at lower cost), ARC-AGI-3 3× next-best, OSWorld 2.0 best-at-any-cost, GDPval-AA v2 SOTA, Zapier AutomationBench ~1.5× next-best pass rate, HLE best cost-efficiency, near-Fable CursorBench (within 0.5% at half cost). Strengths: reasoning on novel problems (ARC-AGI-3), computer use, agentic coding at Opus-level pricing (\(5/\)25). Weaknesses: behind Fable on raw capability; well behind Mythos 5 on offensive cyber (exploit development).

3.3 GPT-5.6 (Sol / Terra / Luna) — the strongest coding & efficiency workhorse

Sol is the benchmark king on agentic coding (AA Coding Index 80, Terminal-Bench 88.8, DeepSWE 72.7), browsing (BrowseComp 92.2 ultra), computer use (OSWorld 62.6), academics (GPQA 94.6, FrontierMath 89), long context (MRCR 91.5), cybersecurity (ExploitBench 73.5, CTF 96.7, SEC-Bench Pro 71.2), and professional knowledge work (Agents' Last Exam 53.6 — 13 points over Fable 5). Terra/Luna extend the family's price-performance: Luna outperforms Fable 5 on Agents' Last Exam at ~1/16th the cost, and matches year-old frontier models at ~6% of the cost. Weaknesses: ARC-AGI-3 is low (7.78%) — novel abstraction is its blind spot; AutomationBench (18.1) is the worst of the frontier group; ultra's multi-agent parallelism is the only way to its best scores.

3.4 DeepSeek-V4-Flash-0731 — the agentic open-source cost leader

A Flash-class model that leads on terminal/agentic tool use among open models: Terminal-Bench 2.1 (82.7, near GPT-5.6 Sol), Toolathlon verified (70.3 — the top reported figure here), Cybergym (76.7), NL2Repo (54.2), Automation Bench Public (25.1). Strengths: very strong agent harness integration (Responses API, Codex-adapted), low cost. Weaknesses: far behind the frontier on professional knowledge work (Agent Last Exam 25.2 vs 52.7) and full-stack engineering depth (DeepSWE 54.4); it's a Flash-tier model — V4-Pro (still pending) is the intended frontier competitor.

3.5 Kimi K3 — the best open-weight long-horizon model

The first open 3T-class model. Strongest claims on long-horizon agentic work: BrowseComp 90.4 (tied with GPT-5.6 Sol, with 1M context and no context management), DeepSWE 67.3, plus leadership on in-house productivity suites (OfficeQA Pro, SpreadsheetBench 2, MCP Atlas) and open-ended capability (kernel optimization competitive with Fable 5; built a working Triton-like compiler; designed a chip autonomously). Native multimodal + 1M context + open weights (by Jul 27). Weaknesses: officially concedes a UX/capability gap vs Claude Fable 5 and GPT-5.6 Sol; sensitivity to thinking-history in non-Kimi harnesses; excessive proactiveness on ambiguous tasks.

3.6 GLM-5.2 — the strongest open-source coding model

Best open-source performance on standard coding benchmarks: Terminal-Bench 2.1 (81.0 — within a few points of Claude Opus 4.8's 85.0, ahead of Gemini 3.1 Pro) and SWE-bench Pro (62.1, beating GPT-5.6's mid-tier results from the open-source seat). MIT license, 1M context, IndexShare architecture at 2.9× FLOPs savings. Weaknesses: as an open model it trails the closed frontier on aggregate intelligence (AA Index, Agents' Last Exam class evals); its edge is concentrated in coding/long-horizon engineering rather than breadth.


Bottom line

  • Coding/agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open)
  • Novel reasoning (ARC-AGI-3): Claude Opus 5
  • Knowledge work/analytics: Claude Fable 5 & GPT-5.6 Sol (tie, different benchmarks), GPT-5.6 Luna (cost)
  • Long-context + multimodal: Claude Fable 5, Kimi K3 (open), GPT-5.6 Sol
  • Cyber: Claude Mythos 5 (restricted), GPT-5.6 Sol (generally available)
  • Price-performance: Claude Opus 5 (premium tier), GPT-5.6 Luna/Terra, DeepSeek-V4-Flash-0731 (open)

References

  • Claude Fable 5 / Mythos 5 launch: https://www.anthropic.com/news/claude-fable-5-mythos-5 · redeployment: https://www.anthropic.com/news/redeploying-fable-5
  • Claude Opus 5: https://www.anthropic.com/news/claude-opus-5
  • GPT-5.6 launch: https://openai.com/index/gpt-5-6/ · price/performance update: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
  • DeepSeek-V4-Flash-0731 (change log): https://api-docs.deepseek.com/updates · API models: https://api-docs.deepseek.com/
  • Kimi K3 tech blog: https://www.kimi.com/blog/kimi-k3 · weights: https://github.com/MoonshotAI/Kimi-K3 · API: https://platform.kimi.ai/
  • GLM-5.2 release: https://z.ai/blog/glm-5.2 · GitHub: https://github.com/zai-org/GLM-5 · model: https://huggingface.co/zai-org/GLM-5.2
  • Benchmark origin sites referenced in releases: https://www.swebench.com/ · https://www.frontierbench.ai/ · https://agents-last-exam.org/ · https://artificialanalysis.ai/evaluations/gdpval-aa · https://artificialanalysis.ai/ · https://deepswe.datacurve.ai/ · https://www.swe-marathon.org/ · https://www.frontierswe.com/ · https://posttrainbench.com/ · https://www.vals.ai/benchmarks/programbench · https://lastexam.ai/ · https://arcprize.org/arc-agi · https://livecodebench.github.io/