The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Mathematics

LiveBench mathematics subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Mathematics average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
49 models scored Higher is better Published by LiveBench Apache-2.0
Every model we hold a LiveBench Mathematics score for
All 49 of 49 models
# Model Creator LiveBench Mathematics Known Good Index Measured
1 OpenAI: GPT-5.1 💡 OpenAI 94.5 73 8 Oct 2026
2 gemini-2.5-pro-exp-03-25 Google 90.2 — 8 Oct 2026
3 DeepSeek: R1 💡 DeepSeek 80.7 75 8 Oct 2026
4 OpenAI: o1 💡 OpenAI 80.3 69 8 Oct 2026
5 claude-3-7-sonnet-20250219 Anthropic 79.0 74 8 Oct 2026
6 QwQ-32B Alibaba 77.8 69 8 Oct 2026
7 OpenAI: o3 Mini 💡 OpenAI 77.3 86 8 Oct 2026
8 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 75.8 62 8 Oct 2026
9 DeepSeek: DeepSeek V3 0324 DeepSeek 73.5 64 8 Oct 2026
10 gemini-exp-1206 Google DeepMind,Google 72.4 — 8 Oct 2026
11 gemini-2.0-pro-exp-02-05 Google 71.0 74 8 Oct 2026
12 gpt-4.5-preview OpenAI 69.3 54 8 Oct 2026
13 gemini-2.0-flash-001 Google DeepMind,Google 65.6 61 8 Oct 2026
14 o1-mini OpenAI 62.0 65 8 Oct 2026
15 DeepSeek: DeepSeek V3 DeepSeek 60.5 62 8 Oct 2026
16 gemini-2.0-flash-exp Google DeepMind,Google 60.4 — 8 Oct 2026
17 DeepSeek-R1-Distill-Qwen-32B DeepSeek 59.4 53 8 Oct 2026
18 qwen2.5-max Alibaba 58.4 — 8 Oct 2026
19 QwQ-32B-Preview Alibaba 58.3 — 8 Oct 2026
20 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 58.1 63 8 Oct 2026
21 gemini-2.0-flash-lite Google 58.1 — 8 Oct 2026
22 learnlm-1.5-pro-experimental Unknown 57.8 — 8 Oct 2026
23 gemini-2.0-flash-lite-preview-02-05 Google 55.5 — 8 Oct 2026
24 Google: Gemma 3 27B Google 55.4 47 8 Oct 2026
25 grok-2-1212 xAI 54.9 45 8 Oct 2026
26 Dracarys2-72B-Instruct Unknown 54.7 — 8 Oct 2026
27 claude-3-5-sonnet-20241022 Anthropic 52.3 45 8 Oct 2026
28 OpenAI: GPT-4o OpenAI 49.5 27 8 Oct 2026
29 Qwen2.5 Coder 32B Instruct Alibaba 46.6 — 8 Oct 2026
30 claude-3-opus-20240229 Anthropic 43.6 32 8 Oct 2026
31 mistral-large-2411 Mistral 42.5 37 8 Oct 2026
32 Meta: Llama 3.3 70B Instruct Meta 42.2 34 8 Oct 2026
33 Microsoft: Phi 4 Microsoft 42.0 41 8 Oct 2026
34 Perplexity: Sonar Perplexity 41.6 — 8 Oct 2026
35 Dracarys2-Llama-3.1-70B-Instruct Unknown 40.3 — 8 Oct 2026
36 mistral-small-2501 Mistral 39.9 31 8 Oct 2026
37 mistral-small-2503 Mistral 39.4 33 8 Oct 2026
38 amazon.nova-pro-v1:0 Amazon 38.0 — 8 Oct 2026
39 amazon.nova-lite-v1:0 Amazon 36.7 — 8 Oct 2026
40 OpenAI: GPT-4o-mini OpenAI 36.3 31 8 Oct 2026
41 claude-3-5-haiku-20241022 Anthropic 35.5 30 8 Oct 2026
42 amazon.nova-micro-v1:0 Amazon 34.5 — 8 Oct 2026
43 Google: Gemma 2 27B Google 26.5 21 8 Oct 2026
44 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 21.3 — 8 Oct 2026
45 gemma-2-9b-it Google 19.8 13 8 Oct 2026
46 c4ai-command-r-08-2024 Cohere 19.4 — 8 Oct 2026
47 Phi-3-small-8k-instruct Microsoft 17.6 — 8 Oct 2026
48 Phi-3-mini-4k-instruct Microsoft 15.7 — 8 Oct 2026
49 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 13.6 — 8 Oct 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.