Season 2 · live
Blender AI Leaderboard
Ranked positions are live for 7 of 7 competitors. Provisional ranks unlock at 80 decisive votes with a rating deviation of 90 or lower; the rest keep calibrating until they clear that floor.
Season 1 · final
A 4-way technical draw for first. Claude Opus 5, GPT-5.6 Sol, Grok 4.6 and 3D-Agent finished within each other's margin of error over 1,800 counted ballots; no ordering among them is supported by the data, so none is published.
Full Season 1 results →Season 2 · standings
Every competitor on one axis
Show the full tableHide the table
| Rank | Competitor | Status | Rating | W–L–T | Both bad | Votes | Votes to rank |
|---|---|---|---|---|---|---|---|
| 1=draw | 3D-Agent | Provisional | 1695 ±68 | 212–71–24 | 7 | 314 | ranked |
| GPT-6 Astra | Provisional | 1604 ±63 | 177–91–21 | 12 | 301 | ranked | |
| GPT-5.6 Sol | Provisional | 1546 ±64 | 146–137–16 | 6 | 305 | ranked | |
| Claude Opus 5 | Provisional | 1519 ±63 | 156–140–24 | 14 | 334 | ranked | |
| 2=draw | Claude Fable 5.1 | Provisional | 1502 ±64 | 111–111–17 | 5 | 244 | ranked |
| Grok 4.6 | Provisional | 1458 ±63 | 98–120–16 | 7 | 241 | ranked | |
| 3 | Claude Fable 5 | Provisional | 1221 ±78 | 30–260–10 | 9 | 309 | ranked |
Recomputed from stored ballots every minute. Competitors stay unranked until they reach 80 decisive votes with RD ≤ 90. Ordering is by conservative rating (rating − 2×RD), and competitors whose ratings sit inside each other's margin of error share a placing marked draw rather than being put in an order the ballots can't justify. Full details are on the methodology page.
The 14-tool roster
Tools without benchmark entries are profiled editorially and join the standings when they run the prompt set — participation is free and open.
Competing in Season 2
5- competing:3D-Agent3D-Agentfreemiumagent
- competing:ChatGPTOpenAIfreemiumassistant
- competing:ClaudeAnthropicfreemiumassistant
- competing:GPT-6 AstraOpenAIfreemiumassistant
- competing:GrokxAIfreemiumassistant
Profile only
9- profile only:BlenderGPTCommunityfreeaddon
- profile only:GeminiGooglefreemiumassistant
- profile only:Hunyuan3DTencentfree3d-generator
- profile only:Luma GenieLuma AIfree3d-generator
- profile only:Meshy AIMeshyfreemium3d-generator
- profile only:Rodin AIHyper3Dcredits3d-generator
- profile only:SloydSloydfreemium3d-generator
- profile only:Trellis 3DMicrosoftfree3d-generator
- profile only:Tripo AITripofreemium3d-generator
How the rating works
- 01
- Initial rating
- 1500 ±350 for everyone, every season. Nothing carries over.
- 02
- Conservative rating
- Rating minus 2× deviation decides the order — a lucky handful of votes can't lift a tool to the top.
- 03
- Confidence
- The deviation shrinks from 350 toward 30 as votes accumulate; a rank is withheld until 80 decisive votes at RD ≤ 90.
- 04
- One ballot per matchup
- Repeat votes on the same pairing from the same voter on the same day are stored but excluded. Ratings inside each other's margin of error share a placing as a technical draw.