https://t.co/VLv2NNwdIk 4.0 is live – our biggest release yet!
Generations with Gemini 3.7 Flash for free for a limited time. No API key required!
Make anything you want, publish it to the new Gallery, and vote on the prompts you want to see added to the official benchmark :)
Also new:
- Generate from any prompt, save your builds, and publish them to the new Gallery
- Completely redesigned MineBench + a much faster Arena
- Accounts + personal model rankings
- Global rankings now use Bradley–Terry + 95% confidence intervals
(Oh and MineBench is now on the App Store :)
Full release notes: https://t.co/fK6j5EVqPI
We've finished benchmarking GPT 6 Sol + Luna, but Opus 5.5 has yet to return anything but empty responses.
Like previous Anthropic releases (at least since adaptive thinking became mandatory), Opus 5.5 is prone to exhausting its entire 128k output budget on thinking tokens, leaving nothing for the actual response.
Though this really shouldn't be happening at all, it's quite surprising to see the same issue even when thinking effort is reduced from max to xhigh.
SWE-bench and LiveCodeBench both publish their evaluation harnesses. Even ARC-AGI publishes benchmarking code while maintaining private evaluation sets
Open-sourcing the implementation isn't the same as publishing the test answers; transparency allows independent scrutiny of how models are configured, evaluated, and scored
Test-set confidentiality does not require keeping the evaluation methodology itself closed
FWIW, here are some MineBench builds Astra generated on release day and today with the same prompt
One-off before/afters don't tell you much about stochastic models; if there even was a change, it could just as likely be from changes in the harness
(top: release, bottom: today)
@TylerPayne600 In our experience, Astra still keeps a slight tendency to overengineer – at least to the point where we prefer to oversee most design decisions still ^^
GPT-6 Astra Pro is the first model we've tested that seems to be able to handle basically any prompt we've thrown at it.
So we tried something a little ridiculous: we increased MineBench's gridSize from 256³ to 8,192³ and gave Astra this prompt: “A full rendition of New York City, fully accurate, dense, and to scale”
Astra wrote a ~61 KB program that produced 100.1 million blocks. Expanded into JSON, the build weighed 1.65 GB.
That's a 32x increase on every axis, or 32,768x more possible voxel positions. At ~1 meter per block, the canvas spans 8.2 km edge-to-edge.
Here's a preview of Astra's build. Note that, at this scale, our renderer struggles to show the entire build.
More GPT-6 Astra Pro generations from community prompts in the gallery! Astra is the first model on MineBench to accurately represent all continents on a globe.
View more, or add your own prompts and generations: https://t.co/Pc1pT5CKbh
@nicdunz@R2Cdev_ One API call which returned in 31 minutes 42 seconds, so just slightly above the average inference time of ~26minutes we observed in our normal benchmarking set
@ethereaglehq The whole city comes from one program in global coordinates; chunks are just a storage/rendering split
So yes, street layout is preserved across those boundaries ^^
@imoutoftokensFR We'll fine-tune our renderer + workers to handle the heap and memory sizes for these larger builds and then publish them to the gallery for anyone to view and explore ^^
@nicdunz@R2Cdev_ One API call which returned in 31 minutes 42 seconds, so just slightly above the average inference time of ~26minutes we observed in our normal benchmarking set
@BereczMarton This was made using a modified version of our https://t.co/4Clf6QgNgo harness; we're offering free Gemini 3.8 Flash generations if you want to try it yourself :)
Costs for this specific experiment:
https://t.co/2Q82jPwlNX
@Deep_Burner Its raw response was ~38k tokens, which is the tool call our harness then expands into the 1.65 GB JSON build.
With Astra’s API pricing, that’s ~ $1.90 in output tokens. The input tokens are a minute amount and wouldn't add much to the cost ^^
@Deep_Burner Its raw response was ~38k tokens, which is the tool call our harness then expands into the 1.65 GB JSON build.
With Astra’s API pricing, that’s ~ $1.90 in output tokens. The input tokens are a minute amount and wouldn't add much to the cost ^^
@MongeMkt We had to leave a little suspense 😇
We’re further optimizing the renderer so we can properly walk around the full explore view at this scale ^^
GPT-6 Astra Pro is the first model we've tested that seems to be able to handle basically any prompt we've thrown at it.
So we tried something a little ridiculous: we increased MineBench's gridSize from 256³ to 8,192³ and gave Astra this prompt: “A full rendition of New York City, fully accurate, dense, and to scale”
Astra wrote a ~61 KB program that produced 100.1 million blocks. Expanded into JSON, the build weighed 1.65 GB.
That's a 32x increase on every axis, or 32,768x more possible voxel positions. At ~1 meter per block, the canvas spans 8.2 km edge-to-edge.
Here's a preview of Astra's build. Note that, at this scale, our renderer struggles to show the entire build.
that's a really cool idea, but some of the builds for MineBench can exceed the natural world height limits for Minecraft (we're experimenting with some builds that can be as large as thousands of blocks vertically)
even if we could find a way around that, we believe the the storage and hosting costs would be better spent adding more prompts and benchmarking more models :)
This is personally one of our favorite benchmarks!
Inspired by MazeBench, we had GPT-6 Astra build us a maze on MineBench, then tried to solve it ourselves :)
We couldn't make it through, but maybe you can: https://t.co/pupv6CVaui
This is personally one of our favorite benchmarks!
Inspired by MazeBench, we had GPT-6 Astra build us a maze on MineBench, then tried to solve it ourselves :)
We couldn't make it through, but maybe you can: https://t.co/pupv6CVaui