Leaderboard / Model
Grok 4.7
by xAIAssistantGrok apps; API $2 / $6 per million input / output tokens (under 200K-token prompts)Updated 22 Sept 2026
Grok 4.7 is xAI's newest model, released September 21, 2026: a larger base model trained for long, multi-hour tasks, at the same price as Grok 4.6. In Blender it writes Blender Python or drives your scene over MCP — and it opens Season 3 as its own competitor, so its 3D results are judged blind against Claude Opus 5.5, GPT-6 Astra and the rest.
How to use Grok 4.7 with Blender
Paste its generated bpy scripts into Blender's Scripting workspace, or point an MCP-capable client running grok-4.7 at the blender-mcp server for live scene access. The Grok CLI's headless mode makes it easy to script against from a pipeline.
Best for: Procedural and parametric modelling from a script or pipeline at a low per-token price; the blind votes will show how its 3D output compares.
Which AI built it better?
Two random Season 3 entrants answered the same prompt. Pick the better build and the names are revealed. Your vote counts toward the leaderboard.
Picking a matchup…
What Grok 4.7 built
All 12 Season 3 prompts, one script each, exactly as it came out. Click a build to orbit it in 3D, or compare every AI on the same prompt.
Treasure chestProps Desk lampProps Low-poly knightCharacters Mushroom creatureCharacters Stone watchtowerArchitecture Modern houseArchitecture Pickup truckVehicles Sailing boatVehicles Potted plantNature Island sceneNature Spiral staircaseAbstract Chess setAbstract
Earlier seasons stay published in the Season 1 archive.
Arena record
Season 3Ranked
- 1
- Entrants
- 12
- Builds
- of 12 prompts
- 89
- Times voted on
- Grok 4.7Provisional1428 ±6632–52–5
Every entrant starts each season at 1500 ±350 and stays unranked until enough votes are in. Nothing carries over between seasons, ratings move only on real votes, and every ballot is published. How the rating works
What is Grok 4.7?
Grok 4.7 is xAI's model released on September 21, 2026, succeeding Grok 4.6 at the same price. xAI describes it as a larger base model with a longer reinforcement-learning run weighted toward problems that take many hours to complete, and says it works longer and checks its own work more carefully. It has a 500K-token context window, reasoning effort from low to xhigh (default high) and a knowledge cutoff of May 2026.
On the API (model ID grok-4.7) it costs $2 per million input tokens and $6 per million output tokens for prompts under 200K tokens, and $4 / $12 above that — unchanged from Grok 4.6. A faster serving tier, Grok 4.7 Fast, runs at twice the standard rates and is available only in Cursor and Grok Build.
It has no first-party Blender integration: for 3D work it writes Blender Python (bpy) that you run in the Scripting workspace, or drives a live scene through an MCP-capable client connected to the open-source blender-mcp server. It has its own profile, separate from the Grok page, because the arena competes per-model — Grok 4.6 keeps its own Season 1 record.
Grok 4.7 ran the prompt set the day after release at extra-high reasoning effort, writing a single bpy script per prompt executed once in headless Blender 5.1, and entered Season 3 when it opened on September 23. All twelve prompts exported geometry; two scripts (P01 and P03) stopped part-way on their own errors and are published as the partial builds they left rather than retried.
Grok 4.7 at a glance
- Released: September 21, 2026
- API model ID: grok-4.7
- Price: $2 per million input tokens, $6 per million output tokens ($0.50 cached input) for prompts under 200K tokens; $4 / $12 above that
- Grok 4.7 Fast: the same model on faster infrastructure at twice the standard rates, in Cursor and Grok Build only
- Context window: 500K tokens
- Reasoning effort: low, medium, high (default) or xhigh
- Knowledge cutoff: May 2026
Grok 4.7 vs Grok 4.6
In xAI's launch results Grok 4.7 improves on Grok 4.6 across the board at the same price — Terminal-Bench 4.0 37.6% vs 20.3%, CursorBench 4.0 46.3% vs 40.4%, DeepSWE v1.1 71.0% vs 65.2% (at high effort). The same table puts it behind Claude Fable 5.1 on Terminal-Bench and CursorBench.
None of those benchmarks measure 3D. Grok 4.6 has a full arena record: it finished Season 1 inside the four-way technical draw for first, with ten of twelve prompts producing geometry, and those same outputs competed again in Season 2 and return in Season 3. Grok 4.7 ran the same twelve prompts under the same rule — one script, executed once in headless Blender 5.1, no retries, no cleanup — as a separate entry, so the blind votes decide whether the coding gains show up in the geometry.
Arena status
Grok 4.7 ran the twelve prompts on September 22, 2026, the day after its release, at extra-high reasoning effort, and entered Season 3 when it opened on September 23. All twelve prompts exported geometry, and every output is on this page exactly as generated — including P01 and P03, whose scripts stopped part-way on their own errors (a bmesh operator that does not exist, and a matrix-size mismatch). It started at the Glicko-2 default rating, like every other competitor, and stays unranked until it clears the calibration floors.
Strengths and limitations
Strengths
- Strength:Largest measured gain over Grok 4.6 in xAI's launch results is agentic terminal work — Terminal-Bench 4.0 at 37.6% against 20.3% — the skill a scripted or MCP-driven Blender build leans on
- Strength:Same price as Grok 4.6: $2 / $6 per million tokens under 200K-token prompts
- Strength:500K-token context window and adjustable reasoning effort up to xhigh
- Strength:Exported geometry on all twelve arena prompts at extra-high effort, where Grok 4.6 managed ten
Limitations
- Limitation:No blind-vote record yet: it entered Season 3 when it opened on September 23, 2026, so its rating is uncalibrated and nothing here is a quality claim
- Limitation:Not a mesh generator: geometry comes from code, which makes organic and sculpted subjects the structural weak spot
- Limitation:No official Blender connector; integration is via community MCP tooling or pasted scripts
- Limitation:xAI's own results put it behind Claude Fable 5.1 on Terminal-Bench 4.0 and CursorBench 4.0
Grok 4.7 and Blender: FAQ
When was Grok 4.7 released?
xAI released Grok 4.7 on September 21, 2026. It is available in the Grok apps and on the xAI API as grok-4.7.
How much does Grok 4.7 cost?
On the API, $2 per million input tokens and $6 per million output tokens for prompts under 200K tokens, and $4 / $12 above that — the same as Grok 4.6. Grok 4.7 Fast costs twice the standard rates and is only offered in Cursor and Grok Build.
Can Grok 4.7 use Blender?
Yes, through the same two routes as other assistants: it writes Blender Python you run in the Scripting workspace, or an MCP-capable client running it connects to the open-source MCP for Blender server (formerly blender-mcp) and works in a live scene. There is no official xAI Blender connector.
Is Grok 4.7 better than Grok 4.6 at 3D modelling?
Not settled yet. It beats Grok 4.6 on every coding benchmark xAI published, but none of them measure 3D. Both are in Season 3 on the same twelve prompts, and the blind votes will answer it.