The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Language

LiveBench language subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Language average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
49 models scored Higher is better Published by LiveBench Apache-2.0
Every model we hold a LiveBench Language score for
All 49 of 49 models
# Model Creator LiveBench Language Known Good Index Measured
1 OpenAI: GPT-5.1 💡 OpenAI 80.2 73 8 Oct 2026
2 gemini-2.5-pro-exp-03-25 Google 67.8 — 8 Oct 2026
3 OpenAI: o1 💡 OpenAI 65.4 69 8 Oct 2026
4 gpt-4.5-preview OpenAI 61.5 54 8 Oct 2026
5 claude-3-7-sonnet-20250219 Anthropic 59.9 74 8 Oct 2026
6 qwen2.5-max Alibaba 56.3 — 8 Oct 2026
7 claude-3-5-sonnet-20241022 Anthropic 53.8 45 8 Oct 2026
8 QwQ-32B Alibaba 51.4 69 8 Oct 2026
9 gemini-exp-1206 Google DeepMind,Google 51.3 — 8 Oct 2026
10 OpenAI: o3 Mini 💡 OpenAI 50.7 86 8 Oct 2026
11 claude-3-opus-20240229 Anthropic 50.4 32 8 Oct 2026
12 DeepSeek: DeepSeek V3 0324 DeepSeek 49.1 64 8 Oct 2026
13 DeepSeek: R1 💡 DeepSeek 48.5 75 8 Oct 2026
14 OpenAI: GPT-4o OpenAI 47.6 27 8 Oct 2026
15 DeepSeek: DeepSeek V3 DeepSeek 47.5 62 8 Oct 2026
16 grok-2-1212 xAI 45.6 45 8 Oct 2026
17 gemini-2.0-pro-exp-02-05 Google 44.9 74 8 Oct 2026
18 Perplexity: Sonar Perplexity 44.1 — 8 Oct 2026
19 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 42.2 62 8 Oct 2026
20 learnlm-1.5-pro-experimental Unknown 42.0 — 8 Oct 2026
21 o1-mini OpenAI 40.9 65 8 Oct 2026
22 gemini-2.0-flash-001 Google DeepMind,Google 40.7 61 8 Oct 2026
23 mistral-large-2411 Mistral 39.4 37 8 Oct 2026
24 Meta: Llama 3.3 70B Instruct Meta 39.2 34 8 Oct 2026
25 Dracarys2-Llama-3.1-70B-Instruct Unknown 38.8 — 8 Oct 2026
26 gemini-2.0-flash-exp Google DeepMind,Google 38.2 — 8 Oct 2026
27 amazon.nova-pro-v1:0 Amazon 37.0 — 8 Oct 2026
28 claude-3-5-haiku-20241022 Anthropic 35.4 30 8 Oct 2026
29 Google: Gemma 3 27B Google 34.6 47 8 Oct 2026
30 gemini-2.0-flash-lite-preview-02-05 Google 34.3 — 8 Oct 2026
31 Dracarys2-72B-Instruct Unknown 34.1 — 8 Oct 2026
32 gemini-2.0-flash-lite Google 33.6 — 8 Oct 2026
33 Google: Gemma 2 27B Google 32.6 21 8 Oct 2026
34 mistral-small-2501 Mistral 30.5 31 8 Oct 2026
35 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 29.7 — 8 Oct 2026
36 mistral-small-2503 Mistral 29.1 33 8 Oct 2026
37 OpenAI: GPT-4o-mini OpenAI 28.6 31 8 Oct 2026
38 DeepSeek-R1-Distill-Qwen-32B DeepSeek 26.8 53 8 Oct 2026
39 amazon.nova-lite-v1:0 Amazon 25.9 — 8 Oct 2026
40 Microsoft: Phi 4 Microsoft 25.6 41 8 Oct 2026
41 gemma-2-9b-it Google 25.5 13 8 Oct 2026
42 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 23.8 63 8 Oct 2026
43 Qwen2.5 Coder 32B Instruct Alibaba 23.2 — 8 Oct 2026
44 QwQ-32B-Preview Alibaba 21.1 — 8 Oct 2026
45 c4ai-command-r-08-2024 Cohere 16.7 — 8 Oct 2026
46 amazon.nova-micro-v1:0 Amazon 15.8 — 8 Oct 2026
47 Phi-3-small-8k-instruct Microsoft 12.9 — 8 Oct 2026
48 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 11.2 — 8 Oct 2026
49 Phi-3-mini-4k-instruct Microsoft 9.2 — 8 Oct 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.