The Known Good Updated 8 Oct 2026 Subscribe
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

SWE-bench Verified

Real GitHub issues resolved against a verified test harness. Aggregated from Epoch AI's AI Benchmarking Hub (swe_bench_verified.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
32 models scored Higher is better Published by Epoch AI CC-BY 4.0
SWE-bench Verified

SWE-bench Verified as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

⇩
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
Every model we hold a SWE-bench Verified score for
All 32 of 32 models
# Model Creator SWE-bench Verified Known Good Index Measured
1 Anthropic: Claude Opus 4.7 💡 Anthropic 83.5 90 8 Oct 2026
2 gpt-5.5-pre-release OpenAI 80.6 97 8 Oct 2026
3 Google: Gemini 3.5 Flash 💡 Google 79.3 95 8 Oct 2026
4 Anthropic: Claude Opus 4.6 💡 Anthropic 78.7 87 8 Oct 2026
5 Z.ai: GLM 5.2 💡 Z.ai 78.7 91 8 Oct 2026
6 DeepSeek: DeepSeek V4 Pro 0423 💡 DeepSeek 77.6 93 8 Oct 2026
7 Qwen: Qwen3.7 Max 💡 Alibaba 77.3 93 8 Oct 2026
8 OpenAI: GPT-5.4 💡 OpenAI 76.9 89 8 Oct 2026
9 Qwen: Qwen3.6 Max Preview 💡 Alibaba 76.7 89 8 Oct 2026
10 MoonshotAI: Kimi K2.6 💡 Moonshot AI 76.7 92 8 Oct 2026
11 Anthropic: Claude Opus 4.5 💡 Anthropic 76.7 71 8 Oct 2026
12 Google: Gemini 3.1 Pro Preview Custom Tools 💡 Google 75.6 — 8 Oct 2026
13 Google: Gemini 3 Flash Preview 💡 Google 75.4 83 8 Oct 2026
14 Anthropic: Claude Sonnet 4.6 💡 Anthropic 75.2 80 8 Oct 2026
15 OpenAI: GPT-5.3-Codex 💡 OpenAI 74.8 — 8 Oct 2026
16 Z.ai: GLM 5.1 💡 Z.ai 74.2 90 8 Oct 2026
17 MoonshotAI: Kimi K2.5 💡 Moonshot AI 73.8 — 8 Oct 2026
18 OpenAI: GPT-5.2 💡 OpenAI 73.8 79 8 Oct 2026
19 OpenAI: GPT-5 💡 OpenAI 73.6 77 8 Oct 2026
20 Anthropic: Claude Opus 4.1 💡 Anthropic 73.3 49 8 Oct 2026
21 gemini-3-pro-preview Google 72.9 81 8 Oct 2026
22 Z.ai: GLM 5 💡 Z.ai 72.1 77 8 Oct 2026
23 claude-sonnet-4-5-20250929 Anthropic 71.3 66 8 Oct 2026
24 claude-opus-4-20250514 Anthropic 70.7 68 8 Oct 2026
25 OpenAI: GPT-5.1 💡 OpenAI 68.0 73 8 Oct 2026
26 OpenAI: GPT-5 Mini 💡 OpenAI 64.7 66 8 Oct 2026
27 OpenAI: o3 💡 OpenAI 62.3 72 8 Oct 2026
28 claude-3-7-sonnet-20250219 Anthropic 61.0 74 8 Oct 2026
29 Qwen: Qwen3.6 Plus 💡 Alibaba 57.9 79 8 Oct 2026
30 Google: Gemini 2.5 Pro 💡 Google 57.6 65 8 Oct 2026
31 OpenAI: GPT-4.1 OpenAI 48.5 46 8 Oct 2026
32 OpenAI: GPT-4o OpenAI 31.0 27 8 Oct 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the seven components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.