About The Known Good
Reference data and field notes for IT, security, and AI. We aggregate what other people publish about AI models — prices, benchmark results, throughput, human preference ratings — put it on one comparable footing, and show our working. We do not run our own evaluations, and every figure names its source.
A reference site, not a review site. The question it answers is "what does the published evidence say about this model, today, and where did that come from" — not "which model is best".
Everything is aggregated. Benchmark scores come from their publishers, prices and throughput from the marketplaces that serve the models, Elo from arenas that collect human votes. Our contribution is putting them on one scale, matching model identities across sources that disagree about names, and never publishing a number we cannot point back to a source.
Scheduled jobs ingest each source into a staging table, validate every record, sanity-check the row count against the last good run, and only then swap it into the live tables inside one transaction. A source that returns nothing is rejected and the live data is left untouched.
A build step then renders every page you see to flat HTML from a PostgreSQL database. Nothing is computed when you load a page — the interactive bits run in your browser against the same public JSON you can download. There are no accounts, no logins and nothing to sign in to.
and coding result
other people
speed and latency
| Source | Status | Rows written | Last run | Last success | Note |
|---|---|---|---|---|---|
blog |
ok | 20 | 8 Oct 2026 22:05 UTC | 8 Oct 2026 22:05 UTC | feed items=20 valid=20 rejected=0 duplicate_guid=0 | covers=20/20 (NULL cover renders the gradient) | sanity gate: staged 20 rows vs last good 20 | upserted 20 posts (new=0 [note cut at 172 of 186 characters] |
elo |
ok | 1,095 | 8 Oct 2026 04:32 UTC | 8 Oct 2026 04:32 UTC | tts-arena: ranks were 0-based (0..38) and were shifted to 1..39 to match every other arena_kind; staged 1102 rows (image=82 image_edit=59 speech=39 text=413 text_sc=413 video=48 [note cut at 177 of 2041 characters] |
index |
ok | 1,403 | 8 Oct 2026 04:00 UTC | 8 Oct 2026 04:00 UTC | known_good: 65 models qualifying (>=1 each of science, mathematics, coding; of 7 evaluations) | coding: 19 models (>=2 of 3) | math: 77 models (>=2 of 2) | reasoning: 52 models [note cut at 176 of 294 characters] |
pricing |
ok | 469 | 8 Oct 2026 22:01 UTC | 8 Oct 2026 22:02 UTC | openrouter: 469 records, 469 valid, 0 rejected | models.dev: 226 providers, 8463 entries, 0 rejected | models.dev matched 466 of 469 models | staged: 469 models, price 462 [note cut at 171 of 500 characters] |
quality |
ok | 1,070 | 8 Oct 2026 04:00 UTC | 8 Oct 2026 04:00 UTC | attribution: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from https://epoch.ai/benchmarks. Used under CC-BY 4.0. attribution: LiveBench scores [note cut at 171 of 3299 characters] |
What we will not do
- Claim we tested anything. We do not run evaluations. Every chart names the publisher whose measurement it shows.
- Publish cost per task. It needs token counts from an evaluation suite we do not run, so we publish blended price and explain the blend instead. See methodology.
- Take money to move a ranking. The index is equal-weighted, which leaves nothing to negotiate over.
- Guess. A missing figure renders as an em dash, never as a zero and never as an estimate.
- Run accounts. There is a newsletter. There is nothing to log into.
Contact and corrections
Corrections come first. If a figure looks wrong, start at the evaluation page behind it — every score links to the evaluation, and every evaluation links to the publisher's own page, so you can check our number against theirs in two clicks. If ours is the one that is wrong, tell us and we will fix it and say what changed in the changelog.
For anything else — a correction we have not caught, a question about method, or a figure you want that we do not publish yet — the address is [email protected]. One person reads it.
The project also writes up what it finds at Field Notes, which is the best place to reach whoever is arguing about a number this week.
Newsletter and privacy
The newsletter is the only thing we ask for an address for: a daily brief on AI — new models, price moves, security incidents, and the occasional argument. It is double opt-in — you confirm by email before we ever send anything — delivered through a third-party email provider, and your address is used for nothing but sending it. Opens and clicks are counted in aggregate with the identifying information stripped, so we can see how many people opened an issue but not that you did — the privacy policy has the detail. No accounts, no advertising, no third-party analytics tracking you across the site.
Provenance, in one line
Prices and provider telemetry from OpenRouter. Metadata from models.dev. Benchmark results from their publishers, including Epoch AI under CC-BY with the credit that licence requires. Elo from Arena and TTS Arena. Everything listed, licence by licence, on attribution, and downloadable in full from the data page.