The Known Good Updated 29 Jul 2026

Leaderboard

Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.

62 ranked of 947 tracked models All models → Index methodology → Download as CSV →
Known Good Index leaderboard

Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.

# Model Creator Known Good Index Coding Blended $/1M Output tok/s Context
1 gpt-5.5-pre-release OpenAI 97.9 based on 3 of 7 evaluations
2 Google: Gemini 3.1 Pro Preview 💡 Google 95.9 based on 4 of 7 evaluations $4.50 104 1.05M
3 Google: Gemini 3.5 Flash 💡 Google 95.2 based on 3 of 7 evaluations $3.38 126 1.05M
4 Anthropic: Claude Opus 4.7 💡 Anthropic 93.8 based on 5 of 7 evaluations 100.0 $10.00 56 1.00M
5 Qwen: Qwen3.7 Max 💡 Alibaba 93.2 based on 3 of 7 evaluations $1.88 43 1.00M
6 DeepSeek: DeepSeek V4 Pro 💡 DeepSeek 93.2 based on 3 of 7 evaluations $0.54 44 1.05M
7 MoonshotAI: Kimi K2.6 💡 Moonshot AI 92.8 based on 3 of 7 evaluations $1.16 59 262k
8 Z.ai: GLM 5.2 💡 Z.ai 91.3 based on 3 of 7 evaluations $1.16 48 1.05M
9 Qwen: Qwen3.6 Max Preview 💡 Alibaba 90.4 based on 3 of 7 evaluations $2.34 48 262k
10 OpenAI: GPT-5.4 💡 OpenAI 90.1 based on 5 of 7 evaluations 88.9 $5.62 68 1.05M
11 Anthropic: Claude Opus 4.6 💡 Anthropic 87.9 based on 5 of 7 evaluations 89.5 $10.00 36 1.00M
12 Z.ai: GLM 5.1 💡 Z.ai 87.8 based on 3 of 7 evaluations $1.48 53 205k
13 OpenAI: o3 Mini 💡 OpenAI 85.7 based on 4 of 7 evaluations $1.93 106 200k
14 Google: Gemini 3 Flash Preview 💡 Google 83.4 based on 4 of 7 evaluations 77.4 $1.12 89 1.05M
15 gemini-3-pro-preview Google 83.2 based on 5 of 7 evaluations 73.6
16 OpenAI: GPT-5.2 💡 OpenAI 80.4 based on 5 of 7 evaluations 76.2 $4.81 55 400k
17 Anthropic: Claude Sonnet 4.6 💡 Anthropic 79.7 based on 4 of 7 evaluations 70.9 $6.00 44 1.00M
18 OpenAI: GPT-5 💡 OpenAI 77.8 based on 6 of 7 evaluations 67.2 $3.44 48 400k
19 Qwen: Qwen3.6 Plus 💡 Alibaba 77.6 based on 3 of 7 evaluations $0.73 52 1.00M
20 Z.ai: GLM 5 💡 Z.ai 76.6 based on 4 of 7 evaluations 67.4 $1.35 36 205k
21 DeepSeek: R1 💡 DeepSeek 74.7 based on 4 of 7 evaluations $1.15 49 164k
22 deepseek-r1 DeepSeek 74.7 based on 4 of 7 evaluations
23 claude-3-7-sonnet-20250219 Anthropic 73.9 based on 5 of 7 evaluations 71.0
24 OpenAI: GPT-5.1 💡 OpenAI 73.9 based on 6 of 7 evaluations 67.9 $3.44 53 400k
25 OpenAI: o3 💡 OpenAI 73.5 based on 5 of 7 evaluations $3.50 71 200k
26 gemini-2.0-pro-exp-02-05 Google 73.5 based on 3 of 7 evaluations
27 claude-opus-4-5-20251101 Anthropic 71.5 based on 5 of 7 evaluations 77.9
28 OpenAI: o1 💡 OpenAI 69.7 based on 5 of 7 evaluations $26.25 62 200k
29 claude-opus-4-20250514 Anthropic 68.2 based on 4 of 7 evaluations
30 Z.ai: GLM 4.7 💡 Z.ai 68.0 based on 3 of 7 evaluations $0.74 60 205k
31 grok-4-0709 xAI 67.4 based on 3 of 7 evaluations
32 OpenAI: GPT-5 Nano 💡 OpenAI 67.1 based on 4 of 7 evaluations $0.14 98 400k
33 OpenAI: GPT-5 Mini 💡 OpenAI 66.8 based on 6 of 7 evaluations 50.2 $0.69 83 400k
34 claude-haiku-4-5-20251001 Anthropic 66.6 based on 4 of 7 evaluations
35 claude-sonnet-4-5-20250929 Anthropic 65.8 based on 6 of 7 evaluations 61.1
36 o1-mini OpenAI 64.3 based on 4 of 7 evaluations
37 Google: Gemini 2.5 Pro 💡 Google 64.2 based on 4 of 7 evaluations 42.1 $3.44 98 1.05M
38 DeepSeek: DeepSeek V3 0324 DeepSeek 63.8 based on 4 of 7 evaluations $0.48 25 164k
39 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 62.2 based on 4 of 7 evaluations $0.80 25 8k
40 DeepSeek: DeepSeek V3 DeepSeek 62.2 based on 4 of 7 evaluations $0.35 19 164k
41 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 62.0 based on 3 of 7 evaluations
42 OpenAI: gpt-oss-120b 💡 OpenAI 61.1 based on 3 of 7 evaluations $0.07 145 131k
43 gemini-2.0-flash-001 Google DeepMind,Google 60.5 based on 4 of 7 evaluations
44 gpt-4.5-preview OpenAI 53.9 based on 5 of 7 evaluations
45 claude-opus-4-1-20250805 Anthropic 49.2 based on 5 of 7 evaluations 60.3
46 Google: Gemma 3 27B Google 46.1 based on 4 of 7 evaluations $0.17 21 262k
47 gemma-3-27b-it Google 46.1 based on 4 of 7 evaluations
48 OpenAI: GPT-4.1 OpenAI 45.6 based on 5 of 7 evaluations $3.50 44 1.05M
49 grok-2-1212 xAI 44.6 based on 4 of 7 evaluations
50 claude-3-5-sonnet-20241022 Anthropic 44.5 based on 4 of 7 evaluations
51 Microsoft: Phi 4 Microsoft 40.9 based on 4 of 7 evaluations $0.09 59 16k
52 mistral-large-2411 Mistral 37.0 based on 4 of 7 evaluations
53 Meta: Llama 3.3 70B Instruct Meta 33.5 based on 4 of 7 evaluations $0.20 44 131k
54 mistral-small-2503 Mistral 32.5 based on 4 of 7 evaluations
55 claude-3-opus-20240229 Anthropic 31.8 based on 4 of 7 evaluations
56 mistral-small-2501 Mistral 30.6 based on 4 of 7 evaluations
57 OpenAI: GPT-4o-mini OpenAI 30.2 based on 4 of 7 evaluations $0.26 29 128k
58 claude-3-5-haiku-20241022 Anthropic 28.9 based on 4 of 7 evaluations
59 OpenAI: GPT-4o OpenAI 26.4 based on 6 of 7 evaluations 27.2 $4.38 42 128k
60 Google: Gemma 2 27B Google 20.6 based on 4 of 7 evaluations $0.65 16 8k
61 gemma-2-27b-it Google 20.6 based on 4 of 7 evaluations
62 gemma-2-9b-it Google 11.9 based on 4 of 7 evaluations
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.