The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

Humanity's Last Exam

Expert-written questions spanning many specialist fields. Aggregated from Epoch AI's AI Benchmarking Hub (hle_external.csv, column 'Accuracy'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
48 models scored Higher is better Published by Epoch AI CC-BY 4.0
Every model we hold a Humanity's Last Exam score for
All 48 of 48 models
# Model Creator Humanity's Last Exam Known Good Index Measured
1 OpenAI: GPT-6 Astra 💡 OpenAI 54.8 — 8 Oct 2026
2 Anthropic: Claude Fable 5.1 💡 Anthropic 46.5 — 8 Oct 2026
3 Google: Gemini 3.1 Pro Preview 💡 Google 46.4 93 8 Oct 2026
4 Google: Gemini 3.8 Flash 💡 Google 44.5 — 8 Oct 2026
5 OpenAI: GPT-5.4 Pro 💡 OpenAI 44.3 — 8 Oct 2026
6 muse-spark Meta 40.6 — 8 Oct 2026
7 gemini-3-pro-preview Google 37.5 81 8 Oct 2026
8 OpenAI: GPT-5.4 💡 OpenAI 36.2 89 8 Oct 2026
9 Anthropic: Claude Opus 4.7 💡 Anthropic 36.2 90 8 Oct 2026
10 Anthropic: Claude Opus 4.6 💡 Anthropic 34.4 87 8 Oct 2026
11 OpenAI: GPT-5 Pro 💡 OpenAI 31.6 — 8 Oct 2026
12 OpenAI: GPT-5.2 💡 OpenAI 27.8 79 8 Oct 2026
13 OpenAI: GPT-5 💡 OpenAI 25.3 77 8 Oct 2026
14 claude-opus-4-5-20251101-thinking Anthropic 25.2 — 8 Oct 2026
15 MoonshotAI: Kimi K2.5 💡 Moonshot AI 24.4 — 8 Oct 2026
16 OpenAI: GPT-5.1 💡 OpenAI 23.7 73 8 Oct 2026
17 Google: Gemini 2.5 Pro Preview 06-05 💡 Google 21.6 — 8 Oct 2026
18 OpenAI: o3 💡 OpenAI 20.3 72 8 Oct 2026
19 OpenAI: GPT-5 Mini 💡 OpenAI 19.4 66 8 Oct 2026
20 Gemini 2.5 Pro Experimental (March 2025) Google 18.2 — 8 Oct 2026
21 OpenAI: o4 Mini 💡 OpenAI 18.1 — 8 Oct 2026
22 Google: Gemini 2.5 Pro Preview 05-06 💡 Google 17.8 — 8 Oct 2026
23 Anthropic: Claude Opus 4.5 💡 Anthropic 14.2 71 8 Oct 2026
24 claude-sonnet-4-5-20250929-thinking Anthropic 13.7 — 8 Oct 2026
25 Google: Gemini 2.5 Flash 💡 Google 12.1 — 8 Oct 2026
26 claude-opus-4-1-20250805-thinking Anthropic 11.5 — 8 Oct 2026
27 Gemini 2.5 Flash Preview (May 2025) Google 11.0 — 8 Oct 2026
28 Anthropic: Claude Opus 4 💡 Anthropic 10.7 — 8 Oct 2026
29 Google: Gemini 3.1 Flash Lite Preview 💡 Google 8.6 — 8 Oct 2026
30 Z.ai: GLM 4.5 💡 Z.ai 8.3 — 8 Oct 2026
31 Z.ai: GLM 4.5 Air 💡 Z.ai 8.1 — 8 Oct 2026
32 OpenAI: o1-pro 💡 OpenAI 8.1 — 8 Oct 2026
33 Claude 3.7 Sonnet (Thinking) Anthropic 8.0 — 8 Oct 2026
34 OpenAI: o1 💡 OpenAI 8.0 69 8 Oct 2026
35 Anthropic: Claude Opus 4.1 💡 Anthropic 7.9 49 8 Oct 2026
36 Anthropic: Claude Sonnet 4 💡 Anthropic 7.8 — 8 Oct 2026
37 claude-sonnet-4-5-20250929 Anthropic 7.5 66 8 Oct 2026
38 gpt-5.1-instant OpenAI 6.8 — 8 Oct 2026
39 Gemini 2.0 Flash Thinking (January 2025) Google DeepMind,Google 6.6 — 8 Oct 2026
40 Meta: Llama 4 Maverick Meta 5.7 — 8 Oct 2026
41 gpt-4.5-preview OpenAI 5.4 54 8 Oct 2026
42 OpenAI: GPT-4.1 OpenAI 5.4 46 8 Oct 2026
43 gemini-1.5-pro-002 Google 4.6 — 8 Oct 2026
44 Mistral: Mistral Medium 3 Mistral 4.5 — 8 Oct 2026
45 Nova Pro Amazon 4.4 — 8 Oct 2026
46 Claude 3.5 Sonnet (October 2024) Anthropic 4.1 — 8 Oct 2026
47 Nova Lite Amazon 3.6 — 8 Oct 2026
48 OpenAI: GPT-4o OpenAI 2.7 27 8 Oct 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the seven components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.