Leaderboard / Season 1

Season 1 final results

Season 1 ran from 2026-08-13 to 2026-09-07: eight competitors, 90 benchmark outputs across 12 prompts, and 1,800 counted blind ballots. Voting is closed and these numbers are frozen — the table below will not change again.

First place

A 4-way technical draw: Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and 3D-Agent

These four are separated by 108 rating points at the widest. The standard error on those gaps is around 98 points, so no pairing among them reaches significance at p < 0.05. The ballots genuinely cannot order them, and an arena that printed 1-2-3-4 anyway would be inventing a result its own data doesn't contain. They share the title.

Why that's a draw and not a ranking

Glicko-2 reports a rating deviation (RD) alongside every rating, and that RD is the standard error of the estimate. The gap between two competitors therefore has a standard error of √(RDa² + RDb²), and a two-sided z-test says whether the gap is real or noise. Here is that test for every pairing in the winning group:

PairingGapStd. errorp-valueVerdict
Claude Opus 5 vs GPT-5.6 Sol14990.888inseparable
Claude Opus 5 vs Grok 4.6921000.356inseparable
Claude Opus 5 vs 3D-Agent108970.265inseparable
GPT-5.6 Sol vs Grok 4.6781000.434inseparable
GPT-5.6 Sol vs 3D-Agent94970.332inseparable
Grok 4.6 vs 3D-Agent16980.870inseparable

The same test is what closes the group: the fifth-placed finisher is separated from the bottom of the winning group at p = 0.020, which is why the draw stops where it does rather than swallowing the whole table.

One caveat worth stating plainly. 3D-Agent reached the top group on 8 of 12 prompts, not the full set. Fewer prompts means a narrower sample of the benchmark, and its Season 1 coverage is the reason it re-ran the complete set for Season 2.

Final standings, all eight competitors

1,800 counted · 145 duplicates excluded · 151 both-bad

RankCompetitorStatusRatingW–L–TBoth badVotes
1=drawClaude Opus 5Provisional1868 ±70315742019428
GPT-5.6 SolProvisional1854 ±70304861822430
Grok 4.6Provisional1776 ±7118861920278
3D-AgentProvisional1760 ±6722081616323
2=drawClaude Fable 5Provisional1536 ±692482472443562
Claude Opus 4.8Provisional1482 ±701643152170570
Claude Sonnet 5Provisional1348 ±701223112152506
Claude Haiku 4.5Calibrating1066 ±9025411760503

1= Claude Opus 5, GPT-5.6 Sol, Grok 4.6, 3D-Agent (technical draw; ratings inside each other's margin of error)

2= Claude Fable 5, Claude Opus 4.8, Claude Sonnet 5 (technical draw; ratings inside each other's margin of error)

Claude Haiku 4.5 finished unranked: it took 443 decisive votes but never pulled its rating deviation under the calibration floor, so Glicko-2 withholds a position. It is listed rather than hidden.

Who won each prompt

Head-to-head wins on each of the twelve prompts. A prompt winner here is simply whoever took the most blind votes on it — on most prompts the margins are small enough that this is a description, not a verdict.

PromptMost winsRecordBallots
P01 Treasure chestGPT-5.6 Sol37–10160
P02 Desk lampClaude Opus 532–2164
P03 Low-poly knightClaude Opus 532–7130
P04 Mushroom creatureClaude Fable 530–12137
P05 Stone watchtower3D-Agent38–13157
P06 Modern houseGPT-5.6 Sol36–3154
P07 Pickup truck3D-Agent30–11153
P08 Sailing boat3D-Agent31–4146
P09 Potted plantClaude Opus 547–2155
P10 Island sceneGPT-5.6 Sol39–5159
P11 Spiral staircaseGPT-5.6 Sol29–8150
P12 Chess setGPT-5.6 Sol24–3135

Check it yourself

Every ballot behind this table is published under CC BY 4.0, and the Season 1 outputs stay online at the paths voters saw them at. Re-run the Glicko-2 maths over the ballots and you should land on exactly these numbers — if you don't, tell us.

Season 1 ballots (CSV) · Season 1 final standings (JSON) · the 12 benchmark prompts · methodology

Season 2 is live, with every rating back at the default and a rebuilt 3D-Agent covering all twelve prompts.

Season 2 standings →