Leaderboard / Season 2
Season 2 final results
Season 2 ran from 7 September 2026 to 23 September 2026: seven competitors, 82 benchmark outputs across 12 prompts, and 1,687 counted blind ballots. Voting is closed and these numbers are frozen. The table below will not change again.
First place
A five-way technical draw
3D-Agent, GPT-6 Astra, Claude Opus 5, Claude Fable 5.1 and Grok 4.6 are separated by 182 rating points at the widest. The standard error on those gaps is around 91 points, so no pairing among them reaches significance at p < 0.05. The ballots can't order them, and printing them in order anyway would invent a result the data doesn't contain. They share the title.
3D-Agent and GPT-6 Astra finished 7 points apart, about 90 ahead of the rest of the group.
- 1=3D-Agent1681 ±68350–116–3312/12 prompts
- 1=GPT-6 Astra1674 ±65300–149–3112/12 prompts
- 1=Claude Opus 51581 ±63256–227–3512/12 prompts
- 1=Claude Fable 5.11532 ±63193–201–2612/12 prompts
- 1=Grok 4.61499 ±64161–185–2410/12 prompts
Why that's a draw and not a ranking
Glicko-2 reports a rating deviation (RD) alongside every rating, and that RD is the standard error of the estimate. The gap between two competitors therefore has a standard error of √(RDa² + RDb²), and a two-sided z-test says whether the gap is real or noise. Here is that test for every pairing in the winning group:
| Pairing | Gap | Std. error | p-value | Verdict |
|---|---|---|---|---|
| 3D-Agent vs GPT-6 Astra | 7 | 94 | 0.941 | Inseparable |
| 3D-Agent vs Claude Opus 5 | 100 | 93 | 0.281 | Inseparable |
| 3D-Agent vs Claude Fable 5.1 | 149 | 93 | 0.108 | Inseparable |
| 3D-Agent vs Grok 4.6 | 182 | 93 | 0.051 | Inseparable |
| GPT-6 Astra vs Claude Opus 5 | 93 | 91 | 0.304 | Inseparable |
| GPT-6 Astra vs Claude Fable 5.1 | 142 | 91 | 0.117 | Inseparable |
| GPT-6 Astra vs Grok 4.6 | 175 | 91 | 0.055 | Inseparable |
| Claude Opus 5 vs Claude Fable 5.1 | 49 | 89 | 0.582 | Inseparable |
| Claude Opus 5 vs Grok 4.6 | 82 | 90 | 0.361 | Inseparable |
| Claude Fable 5.1 vs Grok 4.6 | 33 | 90 | 0.713 | Inseparable |
The same test is what closes the group. A competitor joins a draw only if it is inseparable from every member, and the next finisher, GPT-5.6 Sol, is separated from 3D-Agent at p = 0.019. That is why the draw stops where it does rather than swallowing the whole table.
Four of the seven entrants competed on the files they submitted in Season 1, unchanged: Claude Opus 5, GPT-5.6 Sol, Claude Fable 5 and Grok 4.6, which covered 10 of the 12 prompts.
Final standings, all seven competitors
1,687 counted · 60 duplicates excluded · 49 both-bad
Show the full tableHide the table
| Rank | Competitor | Status | Rating | W–L–T | Both bad | Votes |
|---|---|---|---|---|---|---|
| 1=Draw | 3D-Agent | Provisional | 1681 ±68 | 350–116–33 | 12 | 511 |
| GPT-6 Astra | Provisional | 1674 ±65 | 300–149–31 | 19 | 499 | |
| Claude Opus 5 | Provisional | 1581 ±63 | 256–227–35 | 18 | 536 | |
| Claude Fable 5.1 | Provisional | 1532 ±63 | 193–201–26 | 7 | 427 | |
| Grok 4.6 | Provisional | 1499 ±64 | 161–185–24 | 13 | 383 | |
| 2 | GPT-5.6 Sol | Provisional | 1462 ±64 | 239–240–25 | 14 | 518 |
| 3 | Claude Fable 5 | Provisional | 1091 ±82 | 46–427–12 | 15 | 500 |
1= — 3D-Agent, GPT-6 Astra, Claude Opus 5, Claude Fable 5.1, Grok 4.6 (technical draw; ratings inside each other's margin of error)
2 — GPT-5.6 Sol
3 — Claude Fable 5
Who won each prompt
Head-to-head wins on each of the twelve prompts. A prompt winner here is whoever took the most blind votes on it. On most prompts the margins are small enough that this is a description, not a verdict.
| Prompt | Most wins | Record | Ballots |
|---|---|---|---|
| P01Treasure chest | 3D-Agent | 36–10 | 130 |
| P02Desk lamp | GPT-6 Astra | 27–12 | 144 |
| P03Low-poly knight | 3D-Agent | 34–9 | 155 |
| P04Mushroom creature | 3D-Agent | 36–10 | 153 |
| P05Stone watchtower | 3D-Agent | 36–8 | 141 |
| P06Modern house | Claude Fable 5.1 | 26–6 | 126 |
| P07Pickup truck | 3D-Agent | 37–5 | 147 |
| P08Sailing boat | GPT-6 Astra | 28–10 | 128 |
| P09Potted plant | Claude Opus 5 | 45–7 | 155 |
| P10Island scene | GPT-6 Astra | 36–9 | 145 |
| P11Spiral staircase | GPT-6 Astra | 29–4 | 119 |
| P12Chess set | Claude Opus 5 | 30–15 | 144 |
Check it yourself. Every ballot behind this table is published under CC BY 4.0, and the Season 2 outputs stay online at the paths voters saw them at. Re-run the Glicko-2 maths over the ballots and you should land on exactly these numbers. If you don't, tell us.
Season 3 is live. Nine competitors, ratings back to zero, and first runs from Claude Opus 5.5 and Grok 4.7.
Season 3 standings →