September 7, 2026

Season 1 ends in a four-way tie — and Season 2 opens

1,800 blind ballots were not enough to separate the top four. Here is why we're publishing that instead of a winner.

The result

Season 1 of the Blender AI Arena is closed. Eight competitors ran the same 12 standardized prompts, one generation each, no retries and no cleanup, and 1,800 counted head-to-head ballots were cast on the results with the tool names hidden.

First place is a four-way technical draw between Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and 3D-Agent. They are not tied on points — they are tied in the only sense that matters, which is that the gaps between them are smaller than the uncertainty on those gaps.

Why we aren't just printing the order

Every rating on this site comes with a rating deviation, and that deviation is the standard error of the estimate. So the gap between two competitors carries its own standard error: √(RDa² + RDb²). Run the test on the widest gap in the top group — Claude Opus 5 over 3D-Agent — and you get 108 rating points against a standard error of 97, for p = 0.265. That is not a lead. That is noise with a direction.

A leaderboard will happily render that as 1, 2, 3, 4 anyway, because tables have rows and rows have an order. The reader then takes away “model X beat model Y” — a claim the underlying votes never made. We would rather publish a less quotable result than a wrong one, so competitors inside each other's margin now share a placing across the whole site, marked as a technical draw.

The draw stops where the statistics stop it. The fifth-placed finisher, Claude Fable 5, is separated from the bottom of the winning group at p = 0.020. The top group is a real finding, not a shrug — it says four quite different systems converged on output quality that 1,800 human comparisons could not tell apart.

What the top four actually have in common: nothing

This is the part worth sitting with. The four systems that tied took different routes to get there. Three of them wrote a single Blender Python script from the prompt text, executed once headlessly. One of them — 3D-Agent — worked inside a live Blender scene through MCP, building real objects with modifiers and materials the way an artist would.

Different vendors, different reasoning budgets, radically different execution models, statistically identical output quality. If you were choosing a tool on Season 1 evidence alone, the honest advice is that the choice should turn on workflow fit, cost, and editability — not on which one is “best at 3D”, because Season 1 says none of them is.

One caveat we'd rather state than bury: 3D-Agent reached the top group having submitted 8 of the 12 prompts, a narrower slice of the benchmark than the others covered.

Season 2 is open

Season 2 starts now, and every rating restarts at the Glicko-2 default. Nothing carries over: a Season 1 co-champion begins on 1500 like a first-time entrant, and the Season 2 table is decided only by Season 2 ballots.

  • 3D-Agent returns rebuilt. It has re-run the complete prompt set on its current release, with updated models underneath — 12 of 12 this time, against 8 in Season 1. Its Season 1 files stay published, in place, as the record of what Season 1 voters actually judged.
  • GPT-6 Astra enters for the first time, on the full prompt set, writing one-shot bpy scripts at extra-high reasoning effort.
  • Claude Haiku 4.5, Claude Sonnet 5 and Claude Opus 4.8 retire from the active roster. Their Season 1 record stays in the archive.
  • The returning entrants — Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and Claude Fable 5 — re-enter on the same prompt set with the artifacts they already submitted.

Season 1's ballots, standings and outputs all stay online permanently. Closing a season freezes it; it never deletes it.

Check the numbers

Every ballot behind the Season 1 result is published under CC BY 4.0, and the per-pairing significance maths is laid out in full on the results page. Re-run Glicko-2 over the raw ballots and you should land on exactly the published table.