@OpenAI's GPT-6 Astra hedges ("maybe", "might", "it seems") about 20x as often as @AnthropicAI's Claude Opus 5.5. Opus usually weighs a few options and commits early, in about 4 out of 5 summaries as opposed to Astra's 1 in 4.
We read 324 of their thinking summaries on Design Arena to see how each one actually goes about building a game. Astra works more like a designer, Opus more like a builder.
Full video in the replies.
We were surprised to find that 50% of Open AI's GPT-6 Astra's driving games have mistakenly reversed steering mechanics.
So what makes users prefer games, and where do the models fall short? We analyzed 4,300+ tournaments in DesignArena to find out.
Beyond looks, there are other interesting axes to explore: whether all scene objects obey the same physics, whether the controls work the way players expect, and whether the world interactions hold up once you start moving.
Full video in the thread. Try it for free on DesignArena!
HOLY SHIT, @fal so good even our agents are impressed.
We’re working on a secret new arena coming soon. Here’s what one of our agents had to say after using @fal H3 Max for video generation.
Muse Spark 1.3 is such an underdog for web design—outstanding websites at a fraction of the cost and half the generation time.
Pretty incredible progress from the
@AIatMeta team: 3 models in 3 months, each a significant improvement over the last.
BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362!
This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price.
Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this category on Design Arena.
Note: GPT-6 Astra is still pending final results.
Congrats to the @Meta team!
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
Yesterday I learned a humanoid can run 100m in 8.86 seconds.
Today @fal broke the video generation speed record: 4.7 seconds on average on Design Arena.
It takes longer to read this sentence than to generate the video.
What a time to be alive.
Congrats to the @fal team.
BREAKING: MiniMax H3 Max sets the new Pareto Frontier for video generation, nearly 50x faster than the base model.
This model is post-trained by @fal on @MiniMax_AI H3, and it's in a league of its own: no other Image to Video model on the arena delivers higher preference at a lower generation time.
Its Image-to-Video generation time is just 6.4 seconds, 18x faster than average, and its Text-to-Video generation time is just 4.7 seconds, 24x faster than the average.
Huge congratulations to the @fal team on this release!
I’ve started using Grok 4.6 High Fast on Cursor. The quality–speed–price tradeoff is really starting to get there.
It is also very good on design related tasks.
Grok 4.6 by @SpaceXAI is 4th in 3D Design on Design Arena with an Elo of 1370.
This puts it ahead of Claude Fable 5 by @AnthropicAI and behind Qwen 3.8 Max by @AlibabaGroup.
Grok 4.6 is an 11 rank and 60 Elo jump from SpaceXAI’s previous model, Grok 4.5, and generates designs 62.6% faster.
Congrats to the @SpaceXAI team on this achievement!
Lots of evals flying around these days.
Here’s a short starting guide to the science of benchmarking, so you can tell which numbers actually mean something! https://t.co/7A9qmnePcS
@catherinehyeo Not sure which part you’re unhappy with, but for me it was scheduled workflows starting unpredictably late. My slightly hacky workaround has been GitHub webhooks + aws Lambda.