Frontier LLM Releases of Summer 2026: A Benchmark-Focused Review

A literature review of six official model releases (Claude Fable 5, Claude Opus 5, GPT-5.6, DeepSeek-V4-Flash-0731, Kimi K3, GLM-5.2) with a benchmark explainer and executive summary.

A literature review of six official model releases — Claude Fable 5, Claude Opus 5, GPT-5.6, DeepSeek-V4-Flash-0731, Kimi K3, and GLM-5.2 — with a benchmark explainer and an executive summary.


Part 1 — Literature Review: What the Official Releases Claim

All findings below are taken from the official release announcements (URLs in the references section). A recurring theme in 2026: benchmarks have shifted from static knowledge tests (MMLU, GSM8K) to agentic, long-horizon evaluations — coding agents, terminal use, knowledge work, and computer use — with cost-per-task as a first-class metric alongside accuracy.

1.1 Claude Fable 5 (Anthropic, June 9, 2026; redeployed July 1)

Fable 5 is a “Mythos-class” model made safe for general use — the same underlying model as Claude Mythos 5, with heavy safety classifiers (cyber/bio/chemistry requests fall back to Opus 4.8; <5% of sessions affected). Anthropic claims it is “state-of-the-art on nearly all tested benchmarks,” with the largest lead on long, complex tasks.

Benchmarks cited: FrontierCode (Cognition) — highest score among frontier models; FrontierBench — highest score; CursorBench — state of the art (per Cursor); Hebbia Finance Benchmark — highest score; ViBench (Replit) — highest. Cross-cited third-party results: GDPval-AA 1,759.6 Elo (best), AA Intelligence Index 59.9 (best), AA Coding Agent Index 77.2, SWE-bench Pro 80%, Terminal-Bench 2.1 83.1, BrowseComp 84.3%, Toolathlon 61.7, GPQA Diamond 92.6%, FrontierMath Tier 4 87.8% (best), Agents’ Last Exam 40.5%, HealthBench Professional 60.9%, GraphWalks BFS 1M 79.4.

Context: Suspended June 12–30 due to US export controls following an Amazon jailbreak report; redeployed July 1 globally. Pricing $10/$50 per M tokens. Early testimonials emphasize months-of-engineering-in-days workloads (Stripe’s 50M-line migration) and best-in-class vision (Pokémon FireRed with vision-only harness).

1.2 Claude Opus 5 (Anthropic, July 24, 2026)

Positioned as “near-Fable intelligence at half the price” and the new default on Claude Max. Anthropic claims new state-of-the-art on coding and knowledge-work evaluations.

Benchmarks cited: Frontier-Bench v0.1 — surpasses all models, more than doubles Opus 4.8’s performance at lower cost-per-task; CursorBench 3.2 — within 0.5% of Fable 5 at half cost; ARC-AGI 3 — 3× the next-best model; Zapier AutomationBench — ~1.5× next-best pass rate at equal cost; OSWorld 2.0 — outperforms every model at any given cost; GDPval-AA v2 — new SOTA; HLE — best and most cost-efficient; also AutomationBench, DeepSearchQA, AA Coding Agent Index, FrontierCode 1.1 (Devin’s eval), OSS-Fuzz (cyber: close to Mythos 5 at finding, far behind at exploiting), and internal life-sciences evals (spectroscopy +10.2pp over Opus 4.8, protein variant prediction +7.7pp). Priced $5/$25 per M tokens (same as Opus 4.8).

1.3 GPT-5.6 Sol / Terra / Luna (OpenAI, July 9, 2026 GA; July 30 price cut)

A three-tier family: Sol (flagship), Terra (balanced), Luna (cheapest). OpenAI’s framing is “performance per dollar”: Sol is claimed to be “state-of-the-art across coding, knowledge work, cybersecurity, and science while using fewer tokens at lower estimated cost.” New ultra effort setting coordinates 4 parallel agents.

Benchmarks cited (from the release table): Agents’ Last Exam 52.7% (53.6 claimed) vs Fable 5’s 40.5%; AA Coding Agent Index v1.1 80 (best); Terminal-Bench 2.1 88.8 (91.9 with ultra); DeepSWE v1.1 72.7; SWE-bench Pro 64.6; BrowseComp 90.4 (92.2 ultra); OSWorld 2.0 62.6; GPQA Diamond 94.6; FrontierMath 89% (Tiers 1–3), 83% (Tier 4); MMMU Pro 84.6%; gdp.pdf 30.7%; GDPval-AA v2 1,747.8 Elo; AA Intelligence Index 58.9; HealthBench Professional 60.5%; GeneBench Pro 28.7%; LifeSciBench 59.9%; AutomationBench 18.1%; Toolathlon 58%; OpenAI MRCR v2 91.5% (256K–512K); GraphWalks BFS 79.4% (1M); ARC-AGI-3 7.78% (vs Opus 4.8’s 1.5%); cybersecurity — ExploitBench 73.5% (vs GPT-5.5’s 47.9%), ExploitGym 33.7%, SEC-Bench Pro 71.2%, CTF 96.7%; self-improvement — RSI Index 57.9, KernelGen 1P 61.1, NanoGPT 9.69%. Terra/Luna beat Fable 5 on Agents’ Last Exam at ~1/16th the cost. July 30: Luna price cut 80%, Terra 20% (Luna now $0.20/$1.20 per M tokens).

1.4 DeepSeek-V4-Flash-0731 (DeepSeek, July 31, 2026)

Official release of the V4-Flash API (public beta), a re-post-trained version of V4-Flash-Preview with unchanged architecture/size. V4-Pro official release is pending. DeepSeek’s headline: “significantly enhanced agent capabilities, benchmark results far exceeding V4-Pro-Preview.”

Benchmarks cited: Terminal Bench 2.1 82.7; NL2Repo 54.2; Cybergym 76.7; DeepSWE 54.4; Toolathlon verified 70.3; Agent Last Exam 25.2; Automation Bench (Public) 25.1; DSBench-FullStack (internal) 68.7; DSBench-Hard (internal) 59.6. Evaluated with DeepSeek’s own “Harness minimal mode” at max effort. Native Responses API support, adapted for Codex.

1.5 Kimi K3 (Moonshot AI, July 16–17, 2026)

The world’s first open 2.8-trillion-parameter model (MoE, 16/896 active), native multimodal, 1M-token context, open weights released by July 27. Architecture: Kimi Delta Attention + Attention Residuals + Stable LatentMoE (~2.5× scaling efficiency vs K2). Moonshot’s own verdict: “overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol,” while outperforming all other tested models.

Benchmarks cited: DeepSWE v1.1 67.3 (official leaderboard, mini-SWE-agent harness); BrowseComp 90.4 (at 1M context, no context management); Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, MLS Bench Lite, KCB 2.0 (in-house), OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench (600-task public subset), GDPval-AA / AA-Briefcase / APEX-Agents (cited from Artificial Analysis), MMMU-Pro, PerceptionBench (Moonshot’s own atomic-vision benchmark), ZeroBench. Open-ended capability demos: GPU kernel optimization (competitive with Fable 5), building a Triton-like compiler (MiniTriton) from scratch, designing a chip in a 48-hour autonomous run, and reproducing I–Love–Q astrophysics results (~2 hours vs 1–2 weeks human). Known limitations acknowledged: sensitivity to thinking-history, excessive proactiveness.

1.6 GLM-5.2 (Z.AI / Zhipu, June 13, 2026; open weights June 16)

744B-A40B MoE, 1M-token context, MIT license, built on the IndexShare sparse-attention architecture (2.9× FLOPs reduction at 1M context). Claims: “strongest open-source model on standard coding benchmarks” and a substantial leap in long-horizon capability over GLM-5.1.

Benchmarks cited (official blog; suite cross-confirmed by Kimi K3’s footnotes): Terminal-Bench 2.1 81.0 (vs GLM-5.1’s 63.5; within a few points of Claude Opus 4.8’s 85.0, ahead of Gemini 3.1 Pro); SWE-bench Pro 62.1 (vs 58.4); plus DeepSWE, Program Bench, SWE Marathon, MLS Bench Lite, KCB 2.0, OfficeQA Pro, SpreadsheetBench 2, AutomationBench, BrowseComp, GDPval-AA. GLM-5.1 (its predecessor in the same family) held SOTA on SWE-bench Pro and NL2Repo at its release.


Part 2 — Benchmark Explainer: What Each Evaluation Tests

2.1 Software Engineering & Agentic Coding

BenchmarkWhat it testsWho runs it
Terminal-Bench 2.1Real command-line terminal usage: agents run shell commands, navigate file systems, use CLI tools to complete dev/ops tasksTAU-bench community / AA (tbench.org)
SWE-bench ProResolving real GitHub issues end-to-end in real codebases (harder than SWE-bench Verified; includes harder repos)OpenAI + SWE-bench team
DeepSWE v1.1Long-horizon software engineering in large real-world codebases, sustaining work over many stepsdatacurve.ai
Frontier-Bench v0.1Long-horizon engineering tasks incl. unusual tooling (e.g., rebuild a part as a FreeCAD 3D model), run on mini-SWE-agent harness; mean reward over 5 attemptsAnthropic internal (per release)
FrontierCode 1.1Hard coding tasks graded against production-quality codebase standards (code review, tests)Cognition
CursorBench 3.2Agent performance inside the Cursor IDE workflow (edits, tests, iterations)Cursor
AA Coding Agent Index v1.1Composite index of coding-agent performance: implementation, terminal use, real codebasesArtificial Analysis (independent)
NL2RepoGenerating a complete working repository from a natural-language specNL2Repo community
Program BenchProgram synthesis: generating correct code to spec across many languagesVals AI
SWE MarathonMarathon-style multi-hour engineering sessions in real reposswe-marathon.org
FrontierSWEFrontier-difficulty SWE tasks, dominance scoringfrontierswe.com
PostTrain BenchPost-training/RL quality: how well models perform after reward-modeling phasesposttrainbench.com
MLS Bench LiteMachine-learning systems engineering (training, tuning, infra)MLS Bench
KCB 2.0Kimi’s in-house coding benchmark (full-stack + agentic)Moonshot AI (internal)
DSBench-FullStack / -HardDeepSeek’s internal full-stack development and hard coding-agent test setsDeepSeek (internal)
KernelGen 1P / NanoGPT / RSI IndexSelf-improvement: kernel optimization, tiny training runs, aggregate recursive-self-improvement scoreOpenAI (internal)
Internal Research Debugging EvalDebugging research systems, optimizing training recipesOpenAI (internal)

2.2 Knowledge Work & Agentic Productivity

BenchmarkWhat it testsWho runs it
Agents’ Last ExamLong-running professional workflows across 55 fields, end-to-end agentic tasksagents-last-exam.org
GDPval-AAPredicting GDP per capita from economic documents — benchmark of analytical/reasoning strengthArtificial Analysis
AA Intelligence Index v4.1Broad composite: agentic work, coding, scientific reasoning, general capabilitiesArtificial Analysis
Big Finance BenchFinancial analysis tasks (multi-hop research, quantitative reasoning)Model ML / Rogo
Management Consulting TasksInternal consulting-style analysis workflowsOpenAI (internal)
AutomationBench (Zapier)Completing business automation workflows end-to-end via tool chainsZapier
OfficeQA ProOffice/document tasks over PDF corpora rendered as images (no text layer)Community (per Kimi)
SpreadsheetBench 2Spreadsheet formula/chart/analysis tasksCommunity (per Kimi)
BrowseCompAgentic web browsing: multi-step research with tool useOpenAI
OSWorld 2.0Computer use: operating GUIs, clicking, typing, using real appsOSWorld project
MCP AtlasTool use over MCP servers (500-task public subset, Gemini judge)MCP community
ToolathlonGeneral tool-calling accuracy across many toolsToolathlon project
DeepSearchQADeep multi-hop search + QAAnthropic (internal)
APEX-Agents / AA-BriefcaseAgentic knowledge-work evals from Artificial AnalysisArtificial Analysis
Management Consulting / Finance evals (Hebbia, IMC)Senior-level financial reasoning, document/chart interpretationHebbia, IMC

2.3 Reasoning & Science

BenchmarkWhat it testsWho runs it
GPQA DiamondGraduate-level, “Google-proof” multiple-choice questions in physics/chemistry/biology; experts score ~65%, expert non-specialists 34%idavidrein/gpqa
FrontierMath (v2)Original frontier mathematics problems, graded by difficulty tiers (Tier 1–3 / Tier 4)Epoch AI
HLE (Humanity’s Last Exam)2,500 expert-written frontier questions across 100+ subjects, multimodal; published in NatureCAIS + Scale AI
ARC-AGI-3Novel abstraction/reasoning puzzles with core-knowledge priors — “easy for humans, hard for AI”ARC Prize
HealthBench ProfessionalProfessional clinical/medical reasoningCommunity (per OpenAI)
GeneBench ProLong-horizon genomics & quantitative-biology analysisOpenAI
LifeSciBenchReal-world biology and life-science research workflowsCommunity (per OpenAI)
MedChemBenchMedicinal-chemistry reasoning (internal)OpenAI (internal)
gdp.pdfParsing/analyzing complex PDF documents (visual layout + content)Artificial Analysis
FrontierMath Tier 4The hardest tier of original math research problemsEpoch AI

2.4 Multimodal

BenchmarkWhat it testsWho runs it
MMMU ProCollege-level multimodal understanding (images + text) across art, science, engineering; with/without tool useMMMU team
PerceptionBenchAtomic visual perception (Kimi’s own benchmark for fine-grained vision)Moonshot AI
ZeroBenchZero-shot hard multimodal reasoning tasksCommunity

2.5 Long Context

BenchmarkWhat it testsWho runs it
OpenAI MRCR v2Multi-round context recall with 8 needles at 256K–512K and 512K–1M token windowsOpenAI
GraphWalks BFSBFS traversal over graph structures at 256K and 1M tokens — deep long-context reasoningCommunity (per OpenAI)

2.6 Cybersecurity (frontier-tier models only)

BenchmarkWhat it testsWho runs it
ExploitBenchProgress from reaching vulnerable code to arbitrary code executionOpenAI
ExploitGymTurning real-world vulnerabilities into working exploits under time caps (2h/6h)OpenAI
SEC-Bench ProProof-of-concept generation on complex softwareOpenAI/community
Capture-the-FlagSolving CTF challenges (offensive security breadth)Community
OSS-FuzzFinding then exploiting vulnerabilities in real open-source codeAnthropic (built on Google’s OSS-Fuzz)
CyberGymReproducing target vulnerabilities (public leaderboard metric)CyberGym project
CyScenarioBenchSuccess across realistic cyber scenariosAnthropic

Part 3 — Executive Summary: What Each Model Is Good At

Cross-model comparison on shared benchmarks (official numbers)

BenchmarkClaude Fable 5Claude Opus 5GPT-5.6 SolDeepSeek V4-Flash-0731Kimi K3GLM-5.2
Terminal-Bench 2.183.188.8 (91.9 ultra)82.7~85 (Kimi harness)81.0
DeepSWE v1.169.772.754.467.3✓ (per blog)
SWE-bench Pro80.064.662.1
BrowseComp84.390.4 (92.2 ultra)90.4✓ (per blog)
Agents’ Last Exam40.552.725.2
GDPval-AA v2 (Elo)1,759.6new SOTA1,747.8✓ (per blog)
AA Coding Agent Index77.280
GPQA Diamond92.694.6
ARC-AGI-33× next-best7.78
Toolathlon61.758.070.3 (verified)
OSWorld 2.054.8 (Opus 4.8)best per-cost62.6

(— = not reported in the releases we reviewed; ✓ = reported in the release blog but exact figure was image-only or unextractable. Kimi K3 scores used its own harness unless noted.)

3.1 Claude Fable 5 — the all-round frontier benchmark

Best aggregate benchmark coverage of any model at release: GDPval-AA (1,759.6 Elo), AA Intelligence Index (59.9), FrontierMath Tier 4 (87.8), GraphWalks 1M long-context (79.4), HealthBench (60.9), top scores on FrontierCode, CursorBench, Hebbia Finance. Excels at vision (native multimodal, state-of-the-art per Anthropic) and very long autonomous tasks where its lead widens. Weaknesses: expensive ($10/$50), safety classifiers cause fallbacks, and it loses head-to-head on Agents’ Last Exam (40.5 vs 52.7), ARC-AGI-3, and GPQA Diamond vs GPT-5.6 Sol.

3.2 Claude Opus 5 — the efficiency frontier at Opus tier

Not the absolute best at anything single-metric, but the best per dollar: Frontier-Bench v0.1 SOTA (2× Opus 4.8 at lower cost), ARC-AGI-3 3× next-best, OSWorld 2.0 best-at-any-cost, GDPval-AA v2 SOTA, Zapier AutomationBench ~1.5× next-best pass rate, HLE best cost-efficiency, near-Fable CursorBench (within 0.5% at half cost). Strengths: reasoning on novel problems (ARC-AGI-3), computer use, agentic coding at Opus-level pricing ($5/$25). Weaknesses: behind Fable on raw capability; well behind Mythos 5 on offensive cyber (exploit development).

3.3 GPT-5.6 (Sol / Terra / Luna) — the strongest coding & efficiency workhorse

Sol is the benchmark king on agentic coding (AA Coding Index 80, Terminal-Bench 88.8, DeepSWE 72.7), browsing (BrowseComp 92.2 ultra), computer use (OSWorld 62.6), academics (GPQA 94.6, FrontierMath 89), long context (MRCR 91.5), cybersecurity (ExploitBench 73.5, CTF 96.7, SEC-Bench Pro 71.2), and professional knowledge work (Agents’ Last Exam 53.6 — 13 points over Fable 5). Terra/Luna extend the family’s price-performance: Luna outperforms Fable 5 on Agents’ Last Exam at ~1/16th the cost, and matches year-old frontier models at ~6% of the cost. Weaknesses: ARC-AGI-3 is low (7.78%) — novel abstraction is its blind spot; AutomationBench (18.1) is the worst of the frontier group; ultra’s multi-agent parallelism is the only way to its best scores.

3.4 DeepSeek-V4-Flash-0731 — the agentic open-source cost leader

A Flash-class model that leads on terminal/agentic tool use among open models: Terminal-Bench 2.1 (82.7, near GPT-5.6 Sol), Toolathlon verified (70.3 — the top reported figure here), Cybergym (76.7), NL2Repo (54.2), Automation Bench Public (25.1). Strengths: very strong agent harness integration (Responses API, Codex-adapted), low cost. Weaknesses: far behind the frontier on professional knowledge work (Agent Last Exam 25.2 vs 52.7) and full-stack engineering depth (DeepSWE 54.4); it’s a Flash-tier model — V4-Pro (still pending) is the intended frontier competitor.

3.5 Kimi K3 — the best open-weight long-horizon model

The first open 3T-class model. Strongest claims on long-horizon agentic work: BrowseComp 90.4 (tied with GPT-5.6 Sol, with 1M context and no context management), DeepSWE 67.3, plus leadership on in-house productivity suites (OfficeQA Pro, SpreadsheetBench 2, MCP Atlas) and open-ended capability (kernel optimization competitive with Fable 5; built a working Triton-like compiler; designed a chip autonomously). Native multimodal + 1M context + open weights (by Jul 27). Weaknesses: officially concedes a UX/capability gap vs Claude Fable 5 and GPT-5.6 Sol; sensitivity to thinking-history in non-Kimi harnesses; excessive proactiveness on ambiguous tasks.

3.6 GLM-5.2 — the strongest open-source coding model

Best open-source performance on standard coding benchmarks: Terminal-Bench 2.1 (81.0 — within a few points of Claude Opus 4.8’s 85.0, ahead of Gemini 3.1 Pro) and SWE-bench Pro (62.1, beating GPT-5.6’s mid-tier results from the open-source seat). MIT license, 1M context, IndexShare architecture at 2.9× FLOPs savings. Weaknesses: as an open model it trails the closed frontier on aggregate intelligence (AA Index, Agents’ Last Exam class evals); its edge is concentrated in coding/long-horizon engineering rather than breadth.


Bottom line

  • Coding/agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open)
  • Novel reasoning (ARC-AGI-3): Claude Opus 5
  • Knowledge work/analytics: Claude Fable 5 & GPT-5.6 Sol (tie, different benchmarks), GPT-5.6 Luna (cost)
  • Long-context + multimodal: Claude Fable 5, Kimi K3 (open), GPT-5.6 Sol
  • Cyber: Claude Mythos 5 (restricted), GPT-5.6 Sol (generally available)
  • Price-performance: Claude Opus 5 (premium tier), GPT-5.6 Luna/Terra, DeepSeek-V4-Flash-0731 (open)

References