Leaderboard / Season 2

Season 2 final results

Season 2 ran from 7 September 2026 to 23 September 2026: seven competitors, 82 benchmark outputs across 12 prompts, and 1,687 counted blind ballots. Voting is closed and these numbers are frozen. The table below will not change again.

1,687Ballots countedFrozen 23 September 2026

First place

A five-way technical draw

3D-Agent, GPT-6 Astra, Claude Opus 5, Claude Fable 5.1 and Grok 4.6 are separated by 182 rating points at the widest. The standard error on those gaps is around 91 points, so no pairing among them reaches significance at p < 0.05. The ballots can't order them, and printing them in order anyway would invent a result the data doesn't contain. They share the title.

3D-Agent and GPT-6 Astra finished 7 points apart, about 90 ahead of the rest of the group.

  1. 1=3D-Agent1681 ±68350–116–3312/12 prompts
  2. 1=GPT-6 Astra1674 ±65300–149–3112/12 prompts
  3. 1=Claude Opus 51581 ±63256–227–3512/12 prompts
  4. 1=Claude Fable 5.11532 ±63193–201–2612/12 prompts
  5. 1=Grok 4.61499 ±64161–185–2410/12 prompts

Why that's a draw and not a ranking

Glicko-2 reports a rating deviation (RD) alongside every rating, and that RD is the standard error of the estimate. The gap between two competitors therefore has a standard error of √(RDa² + RDb²), and a two-sided z-test says whether the gap is real or noise. Here is that test for every pairing in the winning group:

PairingGapStd. errorp-valueVerdict
3D-Agent vs GPT-6 Astra7940.941Inseparable
3D-Agent vs Claude Opus 5100930.281Inseparable
3D-Agent vs Claude Fable 5.1149930.108Inseparable
3D-Agent vs Grok 4.6182930.051Inseparable
GPT-6 Astra vs Claude Opus 593910.304Inseparable
GPT-6 Astra vs Claude Fable 5.1142910.117Inseparable
GPT-6 Astra vs Grok 4.6175910.055Inseparable
Claude Opus 5 vs Claude Fable 5.149890.582Inseparable
Claude Opus 5 vs Grok 4.682900.361Inseparable
Claude Fable 5.1 vs Grok 4.633900.713Inseparable

The same test is what closes the group. A competitor joins a draw only if it is inseparable from every member, and the next finisher, GPT-5.6 Sol, is separated from 3D-Agent at p = 0.019. That is why the draw stops where it does rather than swallowing the whole table.

Four of the seven entrants competed on the files they submitted in Season 1, unchanged: Claude Opus 5, GPT-5.6 Sol, Claude Fable 5 and Grok 4.6, which covered 10 of the 12 prompts.
What carried over from Season 1

Final standings, all seven competitors

1,687 counted · 60 duplicates excluded · 49 both-bad

1=3D-AgentDraw
1681 ±68
1674 ±65
1581 ±63
1532 ±63
1499 ±64
1462 ±64
1091 ±82
rating±RD, the rating deviationcalibrating — unranked until 80 decisive votes at RD ≤ 901500, where every competitor startsOrdered by conservative rating (rating − 2×RD); a shared placing is a technical draw.
Show the full table
RankCompetitorStatusRatingW–L–TBoth badVotes
1=Draw3D-AgentProvisional1681 ±68350–116–3312511
GPT-6 AstraProvisional1674 ±65300–149–3119499
Claude Opus 5Provisional1581 ±63256–227–3518536
Claude Fable 5.1Provisional1532 ±63193–201–267427
Grok 4.6Provisional1499 ±64161–185–2413383
2GPT-5.6 SolProvisional1462 ±64239–240–2514518
3Claude Fable 5Provisional1091 ±8246–427–1215500

1= — 3D-Agent, GPT-6 Astra, Claude Opus 5, Claude Fable 5.1, Grok 4.6 (technical draw; ratings inside each other's margin of error)

2 — GPT-5.6 Sol

3 — Claude Fable 5

Who won each prompt

Head-to-head wins on each of the twelve prompts. A prompt winner here is whoever took the most blind votes on it. On most prompts the margins are small enough that this is a description, not a verdict.

PromptMost winsRecordBallots
P01Treasure chest3D-Agent36–10130
P02Desk lampGPT-6 Astra27–12144
P03Low-poly knight3D-Agent34–9155
P04Mushroom creature3D-Agent36–10153
P05Stone watchtower3D-Agent36–8141
P06Modern houseClaude Fable 5.126–6126
P07Pickup truck3D-Agent37–5147
P08Sailing boatGPT-6 Astra28–10128
P09Potted plantClaude Opus 545–7155
P10Island sceneGPT-6 Astra36–9145
P11Spiral staircaseGPT-6 Astra29–4119
P12Chess setClaude Opus 530–15144

Check it yourself. Every ballot behind this table is published under CC BY 4.0, and the Season 2 outputs stay online at the paths voters saw them at. Re-run the Glicko-2 maths over the ballots and you should land on exactly these numbers. If you don't, tell us.

Season 3 is live. Nine competitors, ratings back to zero, and first runs from Claude Opus 5.5 and Grok 4.7.

Season 3 standings →