The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Coding

LiveBench coding subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Coding average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
49 models scored Higher is better Published by LiveBench Apache-2.0
Every model we hold a LiveBench Coding score for
All 49 of 49 models
# Model Creator LiveBench Coding Known Good Index Measured
1 gemini-2.5-pro-exp-03-25 Google 85.9 — 8 Oct 2026
2 OpenAI: o3 Mini 💡 OpenAI 82.7 86 8 Oct 2026
3 gpt-4.5-preview OpenAI 75.2 54 8 Oct 2026
4 claude-3-7-sonnet-20250219 Anthropic 74.5 74 8 Oct 2026
5 OpenAI: GPT-5.1 💡 OpenAI 72.5 73 8 Oct 2026
6 QwQ-32B Alibaba 72.2 69 8 Oct 2026
7 DeepSeek: DeepSeek V3 0324 DeepSeek 70.9 64 8 Oct 2026
8 OpenAI: o1 💡 OpenAI 69.7 69 8 Oct 2026
9 claude-3-5-sonnet-20241022 Anthropic 67.1 45 8 Oct 2026
10 DeepSeek: R1 💡 DeepSeek 66.7 75 8 Oct 2026
11 qwen2.5-max Alibaba 64.4 — 8 Oct 2026
12 gemini-2.0-pro-exp-02-05 Google 63.5 74 8 Oct 2026
13 gemini-exp-1206 Google DeepMind,Google 63.4 — 8 Oct 2026
14 DeepSeek: DeepSeek V3 DeepSeek 61.8 62 8 Oct 2026
15 Dracarys2-72B-Instruct Unknown 58.9 — 8 Oct 2026
16 Qwen2.5 Coder 32B Instruct Alibaba 56.9 — 8 Oct 2026
17 gemini-2.0-flash-exp Google DeepMind,Google 54.4 — 8 Oct 2026
18 gemini-2.0-flash-001 Google DeepMind,Google 53.9 61 8 Oct 2026
19 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 53.5 62 8 Oct 2026
20 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 51.6 63 8 Oct 2026
21 OpenAI: GPT-4o OpenAI 51.4 27 8 Oct 2026
22 claude-3-5-haiku-20241022 Anthropic 51.4 30 8 Oct 2026
23 o1-mini OpenAI 48.0 65 8 Oct 2026
24 mistral-large-2411 Mistral 47.1 37 8 Oct 2026
25 gemini-2.0-flash-lite Google 47.1 — 8 Oct 2026
26 learnlm-1.5-pro-experimental Unknown 46.9 — 8 Oct 2026
27 grok-2-1212 xAI 46.4 45 8 Oct 2026
28 gemini-2.0-flash-lite-preview-02-05 Google 43.8 — 8 Oct 2026
29 OpenAI: GPT-4o-mini OpenAI 43.1 31 8 Oct 2026
30 Google: Gemma 3 27B Google 39.9 47 8 Oct 2026
31 claude-3-opus-20240229 Anthropic 38.6 32 8 Oct 2026
32 amazon.nova-pro-v1:0 Amazon 38.1 — 8 Oct 2026
33 QwQ-32B-Preview Alibaba 37.2 — 8 Oct 2026
34 Meta: Llama 3.3 70B Instruct Meta 36.6 34 8 Oct 2026
35 Dracarys2-Llama-3.1-70B-Instruct Unknown 36.3 — 8 Oct 2026
36 mistral-small-2503 Mistral 36.2 33 8 Oct 2026
37 Google: Gemma 2 27B Google 36.0 21 8 Oct 2026
38 mistral-small-2501 Mistral 35.3 31 8 Oct 2026
39 Perplexity: Sonar Perplexity 35.1 — 8 Oct 2026
40 DeepSeek-R1-Distill-Qwen-32B DeepSeek 33.7 53 8 Oct 2026
41 Microsoft: Phi 4 Microsoft 30.7 41 8 Oct 2026
42 amazon.nova-lite-v1:0 Amazon 27.5 — 8 Oct 2026
43 gemma-2-9b-it Google 22.5 13 8 Oct 2026
44 Phi-3-small-8k-instruct Microsoft 20.3 — 8 Oct 2026
45 amazon.nova-micro-v1:0 Amazon 20.2 — 8 Oct 2026
46 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 19.1 — 8 Oct 2026
47 c4ai-command-r-08-2024 Cohere 17.9 — 8 Oct 2026
48 Phi-3-mini-4k-instruct Microsoft 15.5 — 8 Oct 2026
49 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 10.4 — 8 Oct 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.