The Known Good Updated 8 Oct 2026 Subscribe

Leaderboard

Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.

65 ranked of 1191 tracked models All models → Index methodology → Download as CSV →
Known Good Index leaderboard

Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.

#▲ Model Creator Known Good Index Coding Blended $/1M Output tok/s Context
1 gpt-5.5-pre-release OpenAI 97.5 based on 3 of 7 evaluations — — — —
2 Google: Gemini 3.5 Flash 💡 Google 94.8 based on 3 of 7 evaluations — $3.38 161 1.05M
3 DeepSeek: DeepSeek V4 Pro 0423 💡 DeepSeek 93.3 based on 3 of 7 evaluations — $0.36 47 1.05M
4 Google: Gemini 3.1 Pro Preview 💡 Google 93.1 based on 4 of 7 evaluations — $4.50 109 1.05M
5 OpenAI: GPT-5.5 💡 OpenAI 92.8 based on 3 of 7 evaluations — $11.25 81 1.05M
6 Qwen: Qwen3.7 Max 💡 Alibaba 92.7 based on 3 of 7 evaluations — $2.21 78 1.00M
7 MoonshotAI: Kimi K2.6 💡 Moonshot AI 92.5 based on 3 of 7 evaluations — $0.94 42 262k
8 Z.ai: GLM 5.2 💡 Z.ai 90.9 based on 3 of 7 evaluations — $0.96 83 1.05M
9 Anthropic: Claude Opus 4.7 💡 Anthropic 90.0 based on 5 of 7 evaluations 97.2 $10.00 57 1.00M
10 Z.ai: GLM 5.1 💡 Z.ai 89.6 based on 3 of 7 evaluations — $1.48 46 205k
11 Qwen: Qwen3.6 Max Preview 💡 Alibaba 89.5 based on 3 of 7 evaluations — $2.31 63 262k
12 OpenAI: GPT-5.4 💡 OpenAI 88.6 based on 5 of 7 evaluations 91.9 $5.62 91 1.05M
13 Anthropic: Claude Opus 4.6 💡 Anthropic 86.6 based on 5 of 7 evaluations 92.5 $10.00 37 1.00M
14 OpenAI: o3 Mini 💡 OpenAI 85.7 based on 4 of 7 evaluations — $1.93 135 200k
15 Google: Gemini 3 Flash Preview 💡 Google 82.8 based on 4 of 7 evaluations 71.6 $1.12 104 1.05M
16 gemini-3-pro-preview Google 81.3 based on 5 of 7 evaluations 75.9 — — —
17 Anthropic: Claude Sonnet 4.6 💡 Anthropic 80.5 based on 4 of 7 evaluations 72.9 $6.00 50 1.00M
18 OpenAI: GPT-5.2 💡 OpenAI 79.3 based on 5 of 7 evaluations 78.6 $4.81 63 400k
19 Qwen: Qwen3.6 Plus 💡 Alibaba 78.7 based on 3 of 7 evaluations — $0.73 38 1.00M
20 Z.ai: GLM 5 💡 Z.ai 77.3 based on 4 of 7 evaluations 69.3 $0.93 57 205k
21 OpenAI: GPT-5 💡 OpenAI 76.9 based on 6 of 7 evaluations 69.0 $3.44 71 400k
22 DeepSeek: R1 💡 DeepSeek 74.8 based on 4 of 7 evaluations — $1.15 22 64k
23 claude-3-7-sonnet-20250219 Anthropic 73.9 based on 5 of 7 evaluations 71.0 — — —
24 gemini-2.0-pro-exp-02-05 Google 73.7 based on 3 of 7 evaluations — — — —
25 OpenAI: GPT-5.1 💡 OpenAI 73.1 based on 6 of 7 evaluations 69.0 $3.44 99 400k
26 OpenAI: o3 💡 OpenAI 72.3 based on 5 of 7 evaluations — $3.50 59 200k
27 Anthropic: Claude Opus 4.5 💡 Anthropic 71.5 based on 5 of 7 evaluations 80.2 $10.00 44 200k
28 OpenAI: o1 💡 OpenAI 69.3 based on 5 of 7 evaluations — $26.25 87 200k
29 QwQ-32B Alibaba 68.9 based on 3 of 7 evaluations — — — —
30 Z.ai: GLM 4.7 💡 Z.ai 68.6 based on 3 of 7 evaluations — $1.00 40 205k
31 claude-opus-4-20250514 Anthropic 68.3 based on 4 of 7 evaluations — — — —
32 grok-4-0709 xAI 67.7 based on 3 of 7 evaluations — — — —
33 OpenAI: GPT-5 Nano 💡 OpenAI 67.6 based on 4 of 7 evaluations — $0.14 91 400k
34 claude-haiku-4-5-20251001 Anthropic 67.2 based on 4 of 7 evaluations — — — —
35 OpenAI: GPT-5 Mini 💡 OpenAI 66.2 based on 6 of 7 evaluations 51.4 $0.69 113 400k
36 claude-sonnet-4-5-20250929 Anthropic 65.9 based on 6 of 7 evaluations 62.6 — — —
37 Qwen: Qwen3.6 35B A3B 💡 Alibaba 65.7 based on 3 of 7 evaluations — $0.36 89 262k
38 Google: Gemini 2.5 Pro 💡 Google 64.7 based on 4 of 7 evaluations 43.3 $3.44 98 1.05M
39 o1-mini OpenAI 64.5 based on 4 of 7 evaluations — — — —
40 DeepSeek: DeepSeek V3 0324 DeepSeek 63.9 based on 4 of 7 evaluations — $0.50 28 164k
41 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 62.5 based on 4 of 7 evaluations — $0.80 — 8k
42 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 62.5 based on 3 of 7 evaluations — — — —
43 DeepSeek: DeepSeek V3 DeepSeek 62.3 based on 4 of 7 evaluations — $0.45 25 164k
44 OpenAI: gpt-oss-120b 💡 OpenAI 61.5 based on 3 of 7 evaluations — $0.07 150 131k
45 gemini-2.0-flash-001 Google DeepMind,Google 60.8 based on 4 of 7 evaluations — — — —
46 gpt-4.5-preview OpenAI 53.8 based on 5 of 7 evaluations — — — —
47 DeepSeek-R1-Distill-Qwen-32B DeepSeek 52.6 based on 3 of 7 evaluations — — — —
48 Anthropic: Claude Opus 4.1 💡 Anthropic 49.4 based on 5 of 7 evaluations 61.6 $30.00 14 200k
49 Google: Gemma 3 27B Google 47.0 based on 4 of 7 evaluations — $0.17 27 131k
50 OpenAI: GPT-4.1 OpenAI 45.5 based on 5 of 7 evaluations — $3.50 101 1.05M
51 grok-2-1212 xAI 45.0 based on 4 of 7 evaluations — — — —
52 claude-3-5-sonnet-20241022 Anthropic 44.9 based on 4 of 7 evaluations — — — —
53 Qwen: Qwen3.5-9B 💡 Alibaba 43.5 based on 3 of 7 evaluations — $0.11 51 262k
54 OpenAI: gpt-oss-20b 💡 OpenAI 41.6 based on 3 of 7 evaluations — $0.04 119 131k
55 Microsoft: Phi 4 Microsoft 41.3 based on 4 of 7 evaluations — $0.09 61 16k
56 mistral-large-2411 Mistral 37.4 based on 4 of 7 evaluations — — — —
57 Meta: Llama 3.3 70B Instruct Meta 34.0 based on 4 of 7 evaluations — $0.15 75 131k
58 mistral-small-2503 Mistral 33.0 based on 4 of 7 evaluations — — — —
59 claude-3-opus-20240229 Anthropic 32.4 based on 4 of 7 evaluations — — — —
60 mistral-small-2501 Mistral 31.2 based on 4 of 7 evaluations — — — —
61 OpenAI: GPT-4o-mini OpenAI 30.9 based on 4 of 7 evaluations — $0.26 78 128k
62 claude-3-5-haiku-20241022 Anthropic 29.6 based on 4 of 7 evaluations — — — —
63 OpenAI: GPT-4o OpenAI 26.7 based on 6 of 7 evaluations 27.2 $4.38 67 128k
64 Google: Gemma 2 27B Google 21.4 based on 4 of 7 evaluations — $0.65 48 8k
65 gemma-2-9b-it Google 12.8 based on 4 of 7 evaluations — — — —
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.