The Known Good Updated 8 Oct 2026 Subscribe

Image Arenas

Elo ratings from blind pairwise votes: a human is shown two anonymous outputs for the same prompt and picks the better one. The votes are collected and published by Arena — we ingest the ratings, we do not run the votes and we do not adjust them.

Text Image Video Speech

Image Arenas

Human preference rating from blind pairwise votes, ingested from Arena

Text to Image ArenaImage Editing Arena
Text to Image Arena

Elo from blind pairwise preference votes, collected and published by Arena. Higher is better. We ingest these ratings; we do not run the votes and we do not run our own evaluations.

⇩
The Known Good
Text to Image Arena — full ranking
81 models rated 10,152,680 votes countedRatings captured 8 Oct 2026 Source: Arena
# Model Creator Elo 95% CI Votes Known Good Index
1 gpt-image-2.5-sunburst OpenAI 1,425 ±7 17,752 —
2 gpt-image-2.5-flare OpenAI 1,398 ±7 16,889 —
3 gpt-image-2 (medium) OpenAI 1,383 ±4 93,334 —
5 mai-image-2.6 microsoft-ai 1,332 ±5 25,091 —
6 gemini-nano-banana-2.1 Google 1,328 ±9 5,310 —
7 reve-2.1 reve 1,302 ±8 8,249 —
8 grok-imagine-image-2.0 (low) xAI 1,294 ±6 14,483 —
9 muse-image Meta 1,274 ±5 42,262 —
10 reve-2.0 reve 1,269 ±6 15,773 —
11 gemini-3.1-flash-image (nano-banana-2) [web-search] Google 1,262 ±4 60,515 —
12 seedream-5.0-pro ByteDance 1,256 ±4 94,181 —
13 qwen-image-3.0-pro Alibaba 1,255 ±6 11,440 —
14 mai-image-2.5 microsoft-ai 1,254 ±4 61,265 —
15 gemini-3.1-flash-lite-image (nano-banana-2-lite) Google 1,250 ±6 17,804 —
16 gemini-3-pro-image-2k (nano-banana-pro) Google 1,248 ±3 175,734 —
17 gpt-image-1.5-high-fidelity OpenAI 1,237 ±3 169,927 —
18 gemini-3-pro-image-preview (nano-banana-pro) Google 1,234 ±5 86,464 —
13 grok-imagine-image-quality xAI 1,228 ±4 48,786 —
19 qwen-image-2.1 Alibaba 1,223 ±7 11,441 —
20 ideogram-4.0-quality Ideogram 1,205 ±4 60,591 —
21 qwen-image-2.0-pro-2026-06-22 Alibaba 1,192 ±6 13,096 —
22 uni-1.1-max Luma AI 1,189 ±6 14,072 —
23 uni-1.1 Luma AI 1,184 ±4 46,699 —
24 mai-image-2 microsoft-ai 1,184 ±5 51,475 —
25 grok-imagine-image xAI 1,170 ±3 272,085 —
26 recraft-v4.1-utility-pro Recraft 1,169 ±11 2,617 —
28 grok-imagine-image-pro xAI 1,162 ±4 99,455 —
29 flux-2-max Black Forest Labs 1,162 ±3 123,203 —
30 flux-2-flex Black Forest Labs 1,157 ±3 156,923 —
31 reve-v1.5 reve 1,155 ±4 39,700 —
32 flux-2-pro Black Forest Labs 1,153 ±3 212,359 —
33 Cosmos3-Super-Text2Image NVIDIA 1,152 ±6 18,485 —
34 gemini-2.5-flash-image-preview (nano-banana) Google 1,150 ±2 890,539 —
35 hunyuan-image-3.0 Tencent 1,150 ±3 180,518 —
36 seedream-4.5 ByteDance 1,149 ±3 300,107 —
37 imagen-ultra-4.0-generate-001 Google 1,148 ±4 388,736 —
38 flux-2-dev Black Forest Labs 1,146 ±4 77,988 —
39 seedream-4-2k ByteDance 1,140 ±7 12,562 —
40 seedream-5.0-lite ByteDance 1,138 ±3 129,950 —
41 wan2.6-t2i Alibaba 1,137 ±3 233,292 —
42 recraft-v4.1-pro Recraft 1,133 ±10 2,795 —
43 imagen-4.0-generate-001 Google 1,129 ±3 539,732 —
44 qwen-image-2512 Alibaba 1,125 ±3 98,288 —
45 krea-2-medium krea 1,124 ±4 58,098 —
46 seedream-4-fal ByteDance 1,117 ±7 11,841 —
47 hidream-o1-image hidream 1,116 ±4 66,906 —
48 wan2.5-t2i-preview Alibaba 1,116 ±3 279,229 —
49 gpt-image-1 OpenAI 1,116 ±3 268,453 —
50 recraft-v4 Recraft 1,115 ±3 137,017 —
51 seedream-4-high-res-fal ByteDance 1,113 ±3 182,685 —
52 krea-2-turbo krea 1,113 ±4 51,552 —
53 krea-2-large krea 1,112 ±4 56,888 —
54 gpt-image-1-mini OpenAI 1,110 ±3 171,058 —
55 wan2.7-image-pro Alibaba 1,104 ±5 30,226 —
56 wan2.7-image Alibaba 1,101 ±5 30,630 —
57 recraft-v4.1-flash Recraft 1,100 ±9 3,516 —
58 mai-image-1 microsoft-ai 1,093 ±4 98,410 —
59 z-image-turbo Alibaba 1,084 ±5 24,808 —
60 seedream-3 ByteDance 1,082 ±5 36,885 —
61 flux-1-kontext-max Black Forest Labs 1,074 ±3 65,510 —
62 flux-2-klein-9b Black Forest Labs 1,070 ±3 150,782 —
63 qwen-image-prompt-extend Alibaba 1,060 ±3 716,773 —
64 flux-1-kontext-pro Black Forest Labs 1,059 ±3 332,367 —
65 imagen-3.0-generate-002 Google 1,058 ±3 359,772 —
66 qwen-image Alibaba 1,057 ±3 84,550 —
67 ideogram-v3-quality Ideogram 1,048 ±4 118,538 —
68 photon Luma AI 1,036 ±4 130,160 —
69 p-image Unknown 1,033 ±3 109,762 —
70 flux-2-klein-4b Black Forest Labs 1,030 ±3 152,389 —
71 runway-gen4 Runway 1,024 ±4 54,361 —
72 recraft-v3 Recraft 1,021 ±4 195,627 —
73 flux-1.1-pro Black Forest Labs 1,016 ±4 70,460 —
74 ideogram-v2 Ideogram 1,013 ±4 72,090 —
75 lucid-origin leonardo-ai 1,013 ±3 287,341 —
76 glm-image Z.ai 1,012 ±9 4,854 —
77 gemini-2.0-flash-preview-image-generation Google 975 ±3 257,227 —
78 flux-1-dev-fp8 Black Forest Labs 969 ±4 49,221 —
79 dall-e-3 OpenAI 968 ±4 238,743 —
80 flux-1-kontext-dev Black Forest Labs 941 ±4 216,225 —
81 stable-diffusion-v35-large Unknown 938 ±5 23,396 —
82 bagel ByteDance 898 ±6 12,363 —
The Known Good
Image Editing Arena — full ranking
59 models rated 53,320,511 votes countedRatings captured 8 Oct 2026 Source: Arena
# Model Creator Elo 95% CI Votes Known Good Index
1 gpt-image-2.5-sunburst OpenAI 1,524 ±5 59,825 —
2 gpt-image-2.5-flare OpenAI 1,481 ±5 57,408 —
3 gpt-image-2 (medium) OpenAI 1,462 ±3 305,825 —
5 mai-image-2.6 microsoft-ai 1,428 ±5 51,384 —
6 gemini-nano-banana-2.1 Google 1,428 ±6 12,985 —
7 grok-imagine-image-2.0 (low) xAI 1,425 ±5 50,427 —
3 mai-image-2.6-preview microsoft-ai 1,420 ±8 5,563 —
8 muse-image Meta 1,403 ±4 142,644 —
9 mai-image-2.5 microsoft-ai 1,401 ±4 190,081 —
10 seedream-5.0-pro ByteDance 1,394 ±3 282,865 —
11 grok-imagine-image-quality xAI 1,391 ±6 38,195 —
12 gemini-3-pro-image-2k (nano-banana-pro) Google 1,390 ±3 651,674 —
13 chatgpt-image-latest-high-fidelity (20251216) OpenAI 1,389 ±3 629,464 —
14 gemini-3.1-flash-image (nano-banana-2) [web-search] Google 1,387 ±3 229,782 —
15 gemini-3-pro-image-preview (nano-banana-pro) Google 1,386 ±3 540,775 —
16 reve-2.1 reve 1,374 ±6 20,792 —
17 gpt-image-1.5-high-fidelity OpenAI 1,370 ±3 654,249 —
18 qwen-image-2.1 Alibaba 1,369 ±6 21,941 —
19 reve-2.0 reve 1,359 ±6 19,265 —
20 ideogram-4.5 Ideogram 1,347 ±5 19,318 —
21 uni-1.1-max Luma AI 1,334 ±5 43,446 —
22 grok-imagine-image xAI 1,329 ±2 756,681 —
23 uni-1.1 Luma AI 1,315 ±3 188,834 —
24 gemini-3.1-flash-lite-image (nano-banana-2-lite) Google 1,314 ±5 65,669 —
25 qwen-image-2.0-pro-2026-06-22 Alibaba 1,303 ±5 56,723 —
26 hunyuan-image-3.0-instruct Tencent 1,303 ±3 331,423 —
27 wan2.7-image-pro Alibaba 1,303 ±4 43,974 —
28 seedream-4.5 ByteDance 1,302 ±2 1,294,217 —
29 wan2.7-image Alibaba 1,301 ±4 44,796 —
30 seedream-5.0-lite ByteDance 1,294 ±3 459,744 —
31 gemini-2.5-flash-image-preview (nano-banana) Google 1,293 ±2 11,383,435 —
32 seedream-4-2k ByteDance 1,270 ±7 210,059 —
33 reve-v1.1 reve 1,262 ±3 672,625 —
34 flux-2-max Black Forest Labs 1,262 ±3 375,861 —
35 kling-image-o1 Kuaishou 1,251 ±4 143,639 —
36 flux-2-pro Black Forest Labs 1,245 ±3 637,986 —
37 qwen-image-edit Alibaba 1,242 ±3 2,021,952 —
38 reve-v1 reve 1,235 ±5 395,313 —
39 qwen-image-edit-2511 Alibaba 1,234 ±3 485,078 —
40 wan2.6-image Alibaba 1,231 ±3 677,385 —
41 flux-2-flex Black Forest Labs 1,225 ±3 443,912 —
42 flux-2-klein-9b Black Forest Labs 1,225 ±3 540,742 —
43 flux-2-dev Black Forest Labs 1,225 ±4 254,281 —
44 seedream-4-high-res-fal ByteDance 1,217 ±3 1,229,850 —
45 p-image-edit Unknown 1,211 ±3 359,880 —
46 seedream-4-fal ByteDance 1,210 ±6 152,113 —
47 reve-v1.1-fast reve 1,207 ±3 553,317 —
48 reve-edit-fast reve 1,198 ±4 230,283 —
49 flux-2-klein-4b Black Forest Labs 1,187 ±3 541,304 —
50 wan2.5-i2i-preview Alibaba 1,181 ±3 723,697 —
51 flux-1-kontext-max Black Forest Labs 1,181 ±3 390,852 —
52 flux-1-kontext-pro Black Forest Labs 1,176 ±3 6,404,567 —
53 flux-1-kontext-dev Black Forest Labs 1,149 ±3 3,639,808 —
54 seededit-3.0 ByteDance 1,139 ±3 4,918,008 —
55 gpt-image-1 OpenAI 1,139 ±3 2,869,941 —
56 gpt-image-1-mini OpenAI 1,124 ±3 682,914 —
57 gemini-2.0-flash-preview-image-generation Google 1,081 ±2 4,942,834 —
58 bagel ByteDance 1,027 ±6 13,488 —
59 step1x-edit stepfun 998 ±4 155,418 —
The Known Good
How to read Elo. A rating is only meaningful against the population that produced it: a model's number in one arena cannot be compared to its number in another, and a rating built on a few hundred votes moves far more than one built on tens of thousands — the 95% confidence interval is the honest guide to that. Elo measures which output a person preferred, not whether it was correct; for correctness see the evaluations and the Known Good Index, which does not include Elo. Ratings are ingested from Arena and are published under the source's own terms — see attribution.