New leader alert 🚨 Opus 5.5 takes the BuildingBench crown, while GPT‑6 Astra and Sol stand out on cost efficiency. 🏗️
🥇 Opus 5.5 (max) jumps from Opus 5’s 74.5 to 86.8—the highest score yet, ahead of GPT‑6 Astra (ultra, 84.3) and Fable 5.1 (max, 81.4).
💸 GPT‑6 Astra stays close at a fraction of the cost: $7.05 per building vs. ~$45.47 for Opus 5.5. Just 2.5 points behind, at 84% lower cost.
⚡ GPT‑6 Sol takes a different tradeoff:
• Ultra: 69.8 at ~$1.92/building
• Max: 69.6 at ~$1.89/building
Below GPT‑5.6 Sol’s 73.9, but roughly 75% cheaper than its $7.54/building. Both settings join the benchmark’s cost–quality frontier.
Watch the leaderboard shift 👇 Building comparisons from Opus 5.5, GPT‑6 Astra, and GPT‑6 Sol in the thread!
What do those scores look like in 3D? 👀🏗️
Opus 5.5 (max) vs. GPT‑6 Sol (ultra) vs. GPT‑6 Astra (ultra), side by side.
Median cost per building:
🟣 Opus 5.5: ~$45.47
🟢 Astra: $7.05
🔵 Sol: ~$1.92
What do those scores look like in 3D? 👀🏗️
Opus 5.5 (max) vs. GPT‑6 Sol (ultra) vs. GPT‑6 Astra (ultra), side by side.
Median cost per building:
🟣 Opus 5.5: ~$45.47
🟢 Astra: $7.05
🔵 Sol: ~$1.92
New leader alert 🚨 Opus 5.5 takes the BuildingBench crown, while GPT‑6 Astra and Sol stand out on cost efficiency. 🏗️
🥇 Opus 5.5 (max) jumps from Opus 5’s 74.5 to 86.8—the highest score yet, ahead of GPT‑6 Astra (ultra, 84.3) and Fable 5.1 (max, 81.4).
💸 GPT‑6 Astra stays close at a fraction of the cost: $7.05 per building vs. ~$45.47 for Opus 5.5. Just 2.5 points behind, at 84% lower cost.
⚡ GPT‑6 Sol takes a different tradeoff:
• Ultra: 69.8 at ~$1.92/building
• Max: 69.6 at ~$1.89/building
Below GPT‑5.6 Sol’s 73.9, but roughly 75% cheaper than its $7.54/building. Both settings join the benchmark’s cost–quality frontier.
Watch the leaderboard shift 👇 Building comparisons from Opus 5.5, GPT‑6 Astra, and GPT‑6 Sol in the thread!
Grok’s leap isn’t just in coding—it’s showing up in 3D simulation, too 🏗️
Grok 4.7 (xhigh) scores 0.783 on BuildingBench, taking Grok from #9 to #3! A big jump from 4.6’s 0.695.
It now trails only GPT‑6 Astra (ultra, 0.843) and Fable 5.1 (max, 0.814).
Median cost per building: $11.43, versus Astra’s $7.11 and Fable 5.1’s $33.15. Around 66% cheaper than Fable.
Congrats @SpaceXAI and @ElonMusk! Coding agents are getting better at building the world.
See the leap from Grok 4.6 to 4.7 for yourself🧐
Missing walls, inside-out surfaces, blurry textures… Grok 4.7 fixes much of this, producing more faithful, complete, and detailed buildings.
On BuildingBench, both at xhigh:
➤ Score: 0.696 → 0.783 (+12.5%)
➤ Median cost per building: $10.59 → $11.43 (+8%)
A substantial quality jump for a modest cost increase. It's surely a strong release! @milichab@veggie_eric@elonmusk
Grok’s leap isn’t just in coding—it’s showing up in 3D simulation, too 🏗️
Grok 4.7 (xhigh) scores 0.783 on BuildingBench, taking Grok from #9 to #3! A big jump from 4.6’s 0.695.
It now trails only GPT‑6 Astra (ultra, 0.843) and Fable 5.1 (max, 0.814).
Median cost per building: $11.43, versus Astra’s $7.11 and Fable 5.1’s $33.15. Around 66% cheaper than Fable.
Congrats @SpaceXAI and @ElonMusk! Coding agents are getting better at building the world.
See the leap from Grok 4.6 to 4.7 for yourself🧐
Missing walls, inside-out surfaces, blurry textures… Grok 4.7 fixes much of this, producing more faithful, complete, and detailed buildings.
On BuildingBench, both at xhigh:
➤ Score: 0.696 → 0.783 (+12.5%)
➤ Median cost per building: $10.59 → $11.43 (+8%)
A substantial quality jump for a modest cost increase. It's surely a strong release! @milichab@veggie_eric@elonmusk
See the leap from Grok 4.6 to 4.7 for yourself🧐
Missing walls, inside-out surfaces, blurry textures… Grok 4.7 fixes much of this, producing more faithful, complete, and detailed buildings.
On BuildingBench, both at xhigh:
➤ Score: 0.696 → 0.783 (+12.5%)
➤ Median cost per building: $10.59 → $11.43 (+8%)
A substantial quality jump for a modest cost increase. It's surely a strong release! @milichab@veggie_eric@elonmusk
@ALX23uz The input is just a task instruction and four photographs of the building. We use each model’s native harness, such as Grok Build for Grok, with no additional custom agent harness.
Check out our blog and GitHub for the full setup!
Astra vs. Fable 5.1: tag, you’re it! 🏃🗽
Tag Game, all in real time, in the NYC simulation Astra itself created.
The catch? The world doesn’t pause while they think. Cue some awkward collisions with pedestrians and cars 😂
Frontier agents are stepping into embodied worlds, and a simple game of tag puts spatial awareness, planning, opponent prediction, and timely reactions to the test. Skills they’ll need to thrive beyond the chat or coding window.
Watch Fable 5.1 beat Astra 2-1 👑👀
Astra vs. Fable 5.1: tag, you’re it! 🏃🗽
Tag Game, all in real time, in the NYC simulation Astra itself created.
The catch? The world doesn’t pause while they think. Cue some awkward collisions with pedestrians and cars 😂
Frontier agents are stepping into embodied worlds, and a simple game of tag puts spatial awareness, planning, opponent prediction, and timely reactions to the test. Skills they’ll need to thrive beyond the chat or coding window.
Watch Fable 5.1 beat Astra 2-1 👑👀
New episode of Running Man 🏃
DeepSeek V4.1 Flash vs. Union Alpha—three real-time rounds through simulated NYC.
Final score: 3–0. No spoilers—watch to see who gets swept 😂
Which model should join the battle next? 👉
Astra vs. Fable 5.1: tag, you’re it! 🏃🗽
Tag Game, all in real time, in the NYC simulation Astra itself created.
The catch? The world doesn’t pause while they think. Cue some awkward collisions with pedestrians and cars 😂
Frontier agents are stepping into embodied worlds, and a simple game of tag puts spatial awareness, planning, opponent prediction, and timely reactions to the test. Skills they’ll need to thrive beyond the chat or coding window.
Watch Fable 5.1 beat Astra 2-1 👑👀