The Known Good Updated 29 Jul 2026

Reference data for AI models

Reference data and field notes for IT, security, and AI. Compare every major model on quality, price, and speed — aggregated from public benchmarks, refreshed continuously, traceable to origin.

Browse the leaderboardHow we source data
Highlights
Quality

Known Good Index · Higher is better · Ingested from Epoch AI (CC-BY) and LiveBench

Price

Blended USD per 1M tokens · Lower is better · Published rates via OpenRouter

Find your best-fit model
Tailored recommendations weighted for your priorities across quality, speed, and cost
Compare any two models head to head
Side-by-side quality, pricing, context, and provider availability
Download the raw data
Every figure as CSV or JSON, free, with source attribution — no account needed
Changeloglive
Price change · 29 Jul
Google: Gemma 4 31B price up 28%
Arena update · 29 Jul
Arena Elo refreshed: 1008 ratings across 7 arenas
New model tracked · 29 Jul
gpt-5-nano-2025-08-07_minimal added from Epoch AI benchmark data
New model tracked · 29 Jul
gpt-5.4-2026-03-05_none added from Epoch AI benchmark data
New model tracked · 29 Jul
gpt-5-2025-08-07_minimal added from Epoch AI benchmark data
New model tracked · 29 Jul
gpt-5.2-2025-12-11_none added from Epoch AI benchmark data
New model tracked · 29 Jul
gemini-3.5-flash_minimal added from Epoch AI benchmark data
Benchmark update · 29 Jul
Benchmark refresh: 885 scores across 13 evaluations
Price change · 29 Jul
Google: Gemma 4 26B A4B price up 17%
Price change · 29 Jul
Google: Gemma 4 31B price down 22%
Price change · 29 Jul
Qwen: Qwen3.6 35B A3B price up 14%
Price change · 29 Jul
Meta: Llama 3.3 70B Instruct price up 27%
Jump to section

Quality

Quality of leading AI models, aggregated from published independent evaluations

Known Good IndexCoding IndexAgentic Index
Known Good Index

Known Good Index v1.3 incorporates 7 evaluations: GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench. Normalised 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
Known Good Index v1.3 combines seven public benchmarks, each min-max normalised across tracked models then weighted equally. We do not run these evaluations ourselves — scores are ingested from Epoch AI (CC-BY) and LiveBench. See the Index methodology for weights, normalisation, and per-eval provenance.

Coding Index

Software engineering and agentic terminal ability

Coding IndexSWE-bench VerifiedTerminal-BenchLiveBench CodingAgentic Index
Coding Index

Coding Index, built from SWE-bench Verified, LiveBench Coding and Terminal-Bench, each min-max normalised across tracked models and averaged, scaled 0–100. Ingested from Epoch AI (CC-BY 4.0) and LiveBench (Apache-2.0). We do not run these evaluations.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Image & Video Leaderboards

Elo from blind human preference votes, with 95% confidence intervals

Text to ImageImage EditingText to VideoImage to Video

Speech Leaderboards

Text-to-speech quality from blind listening preference votes

Text to Speech Arena

Elo rating from blind A/B listening tests. Ingested from TTS Arena, updated daily. Higher is better.

The Known Good

Capability Indices

Capability-specific indices built from the relevant evaluations

MathematicsReasoningAgentic
Mathematics Index

Mathematics Index, built from Mock AIME 2024–25 and MATH Level 5, each min-max normalised across tracked models and averaged, scaled 0–100. Ingested from Epoch AI (CC-BY 4.0). We do not run these evaluations.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Quality Breakdown

Individual evaluations behind the index, as measured by their publishers

GPQA DiamondMock AIME 2024–25MATH Level 5SWE-bench VerifiedLiveBenchHumanity's Last ExamTerminal-Bench
GPQA Diamond

Graduate-level science questions written to be Google-proof. Aggregated from Epoch AI's AI Benchmarking Hub (gpqa_diamond.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations. Higher is better.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Arena Elo

Human preference rating from blind pairwise votes

OverallStyle controlled
Text Arena

Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.

The Known Good

Openness Index

Weight availability and licence permissiveness

Openness Index — distribution

How many tracked models sit in each openness band. Scores weight availability, licence permissiveness, and commercial-use restrictions, derived from models.dev and Epoch AI metadata, not from evaluations. Bars are bands, not models, so bar colour carries no creator here.

All 5 bands · 947 models classified
Each bar is an openness band; the value is how many tracked models sit in itBands: 100 open weights, permissive · 90 open weights, permissive, no retrievable repository · 60 open weights, restricted use · 40 open weights, non-commercial · 5 proprietary
The Known Good

Context Window

Maximum input tokens accepted

Context windowMax output tokens
Context Window

Maximum input context, from provider metadata via models.dev and OpenRouter. Higher is better.

The Known Good

Output Tokens

Maximum tokens a model will emit in a single response

Max output tokensContext window
Max Output Tokens

Provider-declared maximum completion length, from provider metadata via models.dev and OpenRouter. Higher allows longer single-pass generation.

The Known Good

Price and Cost

Published API pricing, blended and by token type

Blended priceInput priceOutput priceCache hit priceQuality per dollar
Blended Price

USD per 1M tokens at a 3:1 input:output blend. Live from OpenRouter, cross-checked against models.dev. Lower is better. Excludes 18 models published at no cost (free and promotional tiers).

The Known Good
We publish blended price rather than cost-per-task. Cost-per-task requires token counts from a proprietary evaluation suite we do not run. Blended price is computed from published rates and is fully reproducible — see pricing methodology.

Speed & Latency

Throughput and time-to-first-token, measured across hosting providers

Output speedLatency (TTFT)

API Provider Performance

How inference hosts compare when serving the same model

Output speedLatencyPriceContext served
Output Speed by Provider: gpt-oss-120b

Median tokens per second per host, serving the same model (gpt-oss-120b). Derived from OpenRouter endpoint telemetry. Higher is better.

The Known Good
Field notes

Hugging Face Wasn't the Target. It Was in the Way.

21 Jul · Field note

Delaware County's Cyberattack and the Missing 911 Boundary

20 Jul · Field note

When the Router Exfiltrates Its Own Configuration

18 Jul · Field note

AI Worked Both Sides of the Ledger This Week. The Lesson Isn't "Patch Faster."

17 Jul · Field note

Pennsylvania Says Its Statewide 911 Disruption Wasn't a Cyberattack. The Alternative May Be Harder to Defend Against.

15 Jul · Field note

PamStealer Skips the Process Chains Defenders Watch. Not All of Them.

2 Jul · Field note

Your Agent Framework Is a Pile of API Keys on a Public IP

1 Jul · Field note

There's No Patch for FortiBleed. Public-Sector Networks Are Where That Hurts Most.

27 Jun · Field note

Four Failures, One Excuse: "Patch Faster" Was Never the Answer

23 Jun · Field note

Vibe Coding Isn't the Problem. Not Understanding the Stack Is.

20 Jun · Field note

I Handed Claude Code the Keys. Turns Out I'm Not the Only One Using Them.

16 Jun · Field note

The Circuit Nobody Could Find: How Exact-Match Searching Nearly Cost Me the Audit

12 Jun · Field note

The HTTP/2 Bomb Sat in Plain Sight for a Decade. An AI Just Had to Read the Code.

4 Jun · Field note

The Copilot Meter Didn't Raise the Price. It Showed You the Bill.

3 Jun · Field note

Designing AI for a Teen Discord Server Without Turning It Into a Surveillance Machine

29 May · Field note

WhatsApp Says No One Can Read Your Messages. A Federal Agent Spent 10 Months Disagreeing.

22 May · Field note

The Patch Queue Is the New Vulnerability

18 May · Field note

The Week the Toolchain Became the Kill Chain

17 May · Field note

ShinyHunters Didn't Breach 9,000 Schools. They Breached One Vendor. Your Institution Inherited the Rest.

8 May · Field note

The Floor Has Been Hit: Navigating the 2026 Systems Engineering Realignment

3 May · Field note
Every number is sourced
OpenRouter — pricing & provider speedmodels.dev — metadataEpoch AI — benchmarks (CC-BY)LiveBench — contamination-free evalsArena — text, image & video EloTTS Arena — speech Elo Methodology & attribution →
Get the field notes
New models, price moves, and the occasional argument. One email per meaningful update — no accounts, no tracking.