Fields to capture: Organization, methodology, top-ranked models, key findings
Organization
Top Model(s)
Key Finding
Date
3. API Pricing Snapshot
Search each provider's official pricing page. Capture current input/output/thinking token prices per million tokens. Note any cached prompt or batch discounts.
3.1 Pricing Table (per million tokens, USD)
Provider
Model
Input Price
Output Price
Thinking Price
Cached Discount
Batch Discount
Max Context
Date Sourced
OpenAI
GPT-5
OpenAI
o3
OpenAI
o4-mini
Anthropic
Claude Opus 4
Anthropic
Claude Sonnet 4
Google
Gemini 3 Pro
Google
Gemini 3 Flash
xAI
Grok 3
DeepSeek
DeepSeek-V3
DeepSeek
DeepSeek-R1
3.2 Recent Price Changes (since last scan)
Provider
Model
Old Price
New Price
Change (%)
Date
4. Cost-per-Intelligence Ratio (for snapshot)
Calculate rough value ratio for top models: benchmark_score / cost_per_common_task. A quick heuristic, not rigorous.
Model
Avg Benchmark Score (normalized)
Est. Cost per 1M Output Tokens
Value Ratio (score / cost)
5. Summary Output Template
Agent fills this after completing all scans above.
5.1 Top 10 Models by Arena Elo
Rank
Model
Elo
Release Date
5.2 Top 5 Models by SWE-bench Verified
Rank
Model
Resolved Rate
Date
5.3 New Entrants Since Last Quarter
Model
Provider
Release Date
Significance
5.4 Biggest Surprises / Trends
Trend
Description
Evidence
5.5 Best Value Models (Intelligence per Dollar)
Model
Rationale
5.6 Open-Source vs. Closed-Source Gap Update
Dimension
Closed-Source Leader
Open-Source Leader
Gap
Date
Overall Quality (Arena Elo)
Coding (SWE-bench)
Math (AIME)
Long Context
6. Scan Instructions
Expected volume: ~20 web fetches (LMSys, LiveCodeBench, SWE-bench, and each provider's official pricing page).
Priority order: Benchmark leaderboards first, then pricing pages, then supplementary human evaluations.
Tolerance: Skip 404s; annotate stale data if the latest scan is older than 3 months. Prefer official sources over blog summaries.
Date stamp every data point. Without a date, a benchmark score or price is unreliable.
All content in English. Translate non-English sources.
Cost-per-intelligence ratio (section 4): Normalize benchmark scores across models on a 0–100 scale. Compute rough cost estimate using a standard task profile (e.g., 500 input + 2000 output tokens) for each model. Ratio = normalized_score / cost.
Output: This guide, filled with data, becomes the snapshot report. Alternatively, produce a separate companion file llm-landscape-snapshot-[yyyy-mm-dd].md with all findings.