Leaderboard
Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.
Known Good Index leaderboard
Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.
| #▲ | Model | Creator | Known Good Index | Coding | Blended $/1M | Output tok/s | Context |
|---|---|---|---|---|---|---|---|
| 1 |
|
OpenAI | 97.9 based on 3 of 7 evaluations | — | — | — | — |
| 2 |
|
95.9 based on 4 of 7 evaluations | — | $4.50 | 104 | 1.05M | |
| 3 |
|
95.2 based on 3 of 7 evaluations | — | $3.38 | 126 | 1.05M | |
| 4 |
|
Anthropic | 93.8 based on 5 of 7 evaluations | 100.0 | $10.00 | 56 | 1.00M |
| 5 |
|
Alibaba | 93.2 based on 3 of 7 evaluations | — | $1.88 | 43 | 1.00M |
| 6 |
|
DeepSeek | 93.2 based on 3 of 7 evaluations | — | $0.54 | 44 | 1.05M |
| 7 |
|
Moonshot AI | 92.8 based on 3 of 7 evaluations | — | $1.16 | 59 | 262k |
| 8 |
|
Z.ai | 91.3 based on 3 of 7 evaluations | — | $1.16 | 48 | 1.05M |
| 9 |
|
Alibaba | 90.4 based on 3 of 7 evaluations | — | $2.34 | 48 | 262k |
| 10 |
|
OpenAI | 90.1 based on 5 of 7 evaluations | 88.9 | $5.62 | 68 | 1.05M |
| 11 |
|
Anthropic | 87.9 based on 5 of 7 evaluations | 89.5 | $10.00 | 36 | 1.00M |
| 12 |
|
Z.ai | 87.8 based on 3 of 7 evaluations | — | $1.48 | 53 | 205k |
| 13 |
|
OpenAI | 85.7 based on 4 of 7 evaluations | — | $1.93 | 106 | 200k |
| 14 |
|
83.4 based on 4 of 7 evaluations | 77.4 | $1.12 | 89 | 1.05M | |
| 15 |
|
83.2 based on 5 of 7 evaluations | 73.6 | — | — | — | |
| 16 |
|
OpenAI | 80.4 based on 5 of 7 evaluations | 76.2 | $4.81 | 55 | 400k |
| 17 |
|
Anthropic | 79.7 based on 4 of 7 evaluations | 70.9 | $6.00 | 44 | 1.00M |
| 18 |
|
OpenAI | 77.8 based on 6 of 7 evaluations | 67.2 | $3.44 | 48 | 400k |
| 19 |
|
Alibaba | 77.6 based on 3 of 7 evaluations | — | $0.73 | 52 | 1.00M |
| 20 |
|
Z.ai | 76.6 based on 4 of 7 evaluations | 67.4 | $1.35 | 36 | 205k |
| 21 |
|
DeepSeek | 74.7 based on 4 of 7 evaluations | — | $1.15 | 49 | 164k |
| 22 |
|
DeepSeek | 74.7 based on 4 of 7 evaluations | — | — | — | — |
| 23 |
|
Anthropic | 73.9 based on 5 of 7 evaluations | 71.0 | — | — | — |
| 24 |
|
OpenAI | 73.9 based on 6 of 7 evaluations | 67.9 | $3.44 | 53 | 400k |
| 25 |
|
OpenAI | 73.5 based on 5 of 7 evaluations | — | $3.50 | 71 | 200k |
| 26 |
|
73.5 based on 3 of 7 evaluations | — | — | — | — | |
| 27 |
|
Anthropic | 71.5 based on 5 of 7 evaluations | 77.9 | — | — | — |
| 28 |
|
OpenAI | 69.7 based on 5 of 7 evaluations | — | $26.25 | 62 | 200k |
| 29 |
|
Anthropic | 68.2 based on 4 of 7 evaluations | — | — | — | — |
| 30 |
|
Z.ai | 68.0 based on 3 of 7 evaluations | — | $0.74 | 60 | 205k |
| 31 |
|
xAI | 67.4 based on 3 of 7 evaluations | — | — | — | — |
| 32 |
|
OpenAI | 67.1 based on 4 of 7 evaluations | — | $0.14 | 98 | 400k |
| 33 |
|
OpenAI | 66.8 based on 6 of 7 evaluations | 50.2 | $0.69 | 83 | 400k |
| 34 |
|
Anthropic | 66.6 based on 4 of 7 evaluations | — | — | — | — |
| 35 |
|
Anthropic | 65.8 based on 6 of 7 evaluations | 61.1 | — | — | — |
| 36 |
|
OpenAI | 64.3 based on 4 of 7 evaluations | — | — | — | — |
| 37 |
|
64.2 based on 4 of 7 evaluations | 42.1 | $3.44 | 98 | 1.05M | |
| 38 |
|
DeepSeek | 63.8 based on 4 of 7 evaluations | — | $0.48 | 25 | 164k |
| 39 |
|
DeepSeek | 62.2 based on 4 of 7 evaluations | — | $0.80 | 25 | 8k |
| 40 |
|
DeepSeek | 62.2 based on 4 of 7 evaluations | — | $0.35 | 19 | 164k |
| 41 | GD gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 62.0 based on 3 of 7 evaluations | — | — | — | — |
| 42 |
|
OpenAI | 61.1 based on 3 of 7 evaluations | — | $0.07 | 145 | 131k |
| 43 | GD gemini-2.0-flash-001 | Google DeepMind,Google | 60.5 based on 4 of 7 evaluations | — | — | — | — |
| 44 |
|
OpenAI | 53.9 based on 5 of 7 evaluations | — | — | — | — |
| 45 |
|
Anthropic | 49.2 based on 5 of 7 evaluations | 60.3 | — | — | — |
| 46 |
|
46.1 based on 4 of 7 evaluations | — | $0.17 | 21 | 262k | |
| 47 |
|
46.1 based on 4 of 7 evaluations | — | — | — | — | |
| 48 |
|
OpenAI | 45.6 based on 5 of 7 evaluations | — | $3.50 | 44 | 1.05M |
| 49 |
|
xAI | 44.6 based on 4 of 7 evaluations | — | — | — | — |
| 50 |
|
Anthropic | 44.5 based on 4 of 7 evaluations | — | — | — | — |
| 51 |
|
Microsoft | 40.9 based on 4 of 7 evaluations | — | $0.09 | 59 | 16k |
| 52 |
|
Mistral | 37.0 based on 4 of 7 evaluations | — | — | — | — |
| 53 |
|
Meta | 33.5 based on 4 of 7 evaluations | — | $0.20 | 44 | 131k |
| 54 |
|
Mistral | 32.5 based on 4 of 7 evaluations | — | — | — | — |
| 55 |
|
Anthropic | 31.8 based on 4 of 7 evaluations | — | — | — | — |
| 56 |
|
Mistral | 30.6 based on 4 of 7 evaluations | — | — | — | — |
| 57 |
|
OpenAI | 30.2 based on 4 of 7 evaluations | — | $0.26 | 29 | 128k |
| 58 |
|
Anthropic | 28.9 based on 4 of 7 evaluations | — | — | — | — |
| 59 |
|
OpenAI | 26.4 based on 6 of 7 evaluations | 27.2 | $4.38 | 42 | 128k |
| 60 |
|
20.6 based on 4 of 7 evaluations | — | $0.65 | 16 | 8k | |
| 61 |
|
20.6 based on 4 of 7 evaluations | — | — | — | — | |
| 62 |
|
11.9 based on 4 of 7 evaluations | — | — | — | — |
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.