The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Reasoning

LiveBench reasoning subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Reasoning average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
49 models scored Higher is better Published by LiveBench Apache-2.0
Every model we hold a LiveBench Reasoning score for
All 49 of 49 models
# Model Creator LiveBench Reasoning Known Good Index Measured
1 OpenAI: GPT-5.1 💡 OpenAI 95.8 73 8 Oct 2026
2 OpenAI: o1 💡 OpenAI 91.6 69 8 Oct 2026
3 gemini-2.5-pro-exp-03-25 Google 89.8 — 8 Oct 2026
4 OpenAI: o3 Mini 💡 OpenAI 89.6 86 8 Oct 2026
5 claude-3-7-sonnet-20250219 Anthropic 87.8 74 8 Oct 2026
6 QwQ-32B Alibaba 83.5 69 8 Oct 2026
7 DeepSeek: R1 💡 DeepSeek 83.2 75 8 Oct 2026
8 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 78.2 62 8 Oct 2026
9 o1-mini OpenAI 72.3 65 8 Oct 2026
10 gpt-4.5-preview OpenAI 71.1 54 8 Oct 2026
11 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 67.6 63 8 Oct 2026
12 DeepSeek: DeepSeek V3 0324 DeepSeek 65.8 64 8 Oct 2026
13 gemini-2.0-pro-exp-02-05 Google 60.1 74 8 Oct 2026
14 gemini-2.0-flash-exp Google DeepMind,Google 59.1 — 8 Oct 2026
15 QwQ-32B-Preview Alibaba 57.7 — 8 Oct 2026
16 gemini-exp-1206 Google DeepMind,Google 57.0 — 8 Oct 2026
17 DeepSeek: DeepSeek V3 DeepSeek 56.8 62 8 Oct 2026
18 claude-3-5-sonnet-20241022 Anthropic 56.7 45 8 Oct 2026
19 OpenAI: GPT-4o OpenAI 55.8 27 8 Oct 2026
20 gemini-2.0-flash-001 Google DeepMind,Google 55.2 61 8 Oct 2026
21 grok-2-1212 xAI 54.8 45 8 Oct 2026
22 DeepSeek-R1-Distill-Qwen-32B DeepSeek 52.2 53 8 Oct 2026
23 qwen2.5-max Alibaba 51.4 — 8 Oct 2026
24 Meta: Llama 3.3 70B Instruct Meta 50.8 34 8 Oct 2026
25 gemini-2.0-flash-lite-preview-02-05 Google 50.1 — 8 Oct 2026
26 Microsoft: Phi 4 Microsoft 47.8 41 8 Oct 2026
27 Dracarys2-72B-Instruct Unknown 47.4 — 8 Oct 2026
28 Perplexity: Sonar Perplexity 46.2 — 8 Oct 2026
29 gemini-2.0-flash-lite Google 44.9 — 8 Oct 2026
30 mistral-small-2503 Mistral 44.8 33 8 Oct 2026
31 Dracarys2-Llama-3.1-70B-Instruct Unknown 44.7 — 8 Oct 2026
32 Google: Gemma 3 27B Google 43.8 47 8 Oct 2026
33 mistral-large-2411 Mistral 43.5 37 8 Oct 2026
34 learnlm-1.5-pro-experimental Unknown 43.4 — 8 Oct 2026
35 Qwen2.5 Coder 32B Instruct Alibaba 42.1 — 8 Oct 2026
36 claude-3-opus-20240229 Anthropic 40.6 32 8 Oct 2026
37 amazon.nova-lite-v1:0 Amazon 36.7 — 8 Oct 2026
38 mistral-small-2501 Mistral 36.4 31 8 Oct 2026
39 OpenAI: GPT-4o-mini OpenAI 32.8 31 8 Oct 2026
40 amazon.nova-pro-v1:0 Amazon 32.6 — 8 Oct 2026
41 Google: Gemma 2 27B Google 28.1 21 8 Oct 2026
42 claude-3-5-haiku-20241022 Anthropic 28.1 30 8 Oct 2026
43 Phi-3-mini-4k-instruct Microsoft 26.8 — 8 Oct 2026
44 amazon.nova-micro-v1:0 Amazon 25.1 — 8 Oct 2026
45 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 24.8 — 8 Oct 2026
46 c4ai-command-r-08-2024 Cohere 21.9 — 8 Oct 2026
47 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 16.3 — 8 Oct 2026
48 Phi-3-small-8k-instruct Microsoft 15.9 — 8 Oct 2026
49 gemma-2-9b-it Google 15.2 13 8 Oct 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.