Design Arena is excited to introduce the company behind our work: Intelligence.
Over the past year, millions of people have helped us evaluate how well AI models perform on real-world creative tasks - where quality is subjective, outcomes are difficult to verify, and human judgment matters.
Design Arena is only the beginning. Let us know what you want to see next!
Today, we're introducing @Intelligence_ai.
In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries.
We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities.
Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier.
People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them.
Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence.
We've helped the world's leading frontier labs break the news on their SOTA capabilities.
What's the limit? Join us and find out.
Gemini 3.8 Live by @GoogleDeepMind is now available on Design Arena!
Built for more interactive and capable conversations, this model brings upgraded reasoning, near real-time visual understanding, automatic detection across 97 languages, and background tool calling.
Congrats to the @GoogleDeepMind team on the launch!
We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI.
The models talk, think, and handle tasks in the background without breaking your flow. 🧵
BREAKING: @PrunaAI establishes a new Pareto frontier for Speed vs. Preference on Video Editing Arena!
Its debut video editing model, P‑Video‑Edit, offers both Draft and Standard modes - giving users a faster option for iteration before generating the final-quality edit.
Both configurations are live now. Try them today, and congratulations to the team!
DeepSeek-V4.1-Flash by @deepseek_ai is back in the top 10, but what's crazier is that 6 out of the 7 models on the Price vs Preference Pareto frontier are all open weights
BREAKING: DeepSeek‑V4.1‑Flash takes 6th overall on Design Arena with an Elo of 1347!
This marks a 39-position jump over the next-highest DeepSeek model - and DeepSeek’s return to a top-10 placement on Design Arena.
The model ranks in the same performance band as Claude Fable 5.1 on real-world frontend design tasks.
Congratulations to the @deepseek_ai team!
Open weights are on the verge of sweeping Design Arena’s price–preference frontier.
6 of the 7 models defining it are already open weights. If Meta’s planned Muse Spark open-weights release includes Muse Spark 1.3 Max, all seven would be open weights.
The frontier could soon be entirely open.
DeepSeek-V4.1-Flash by @deepseek_ai is now available on Design Arena!
Built as the smallest model in DeepSeek’s new architecture family, DeepSeek-V4.1-Flash brings native visual understanding with a focus on faster inference, higher throughput, and greater efficiency. The model is designed to deliver stronger capabilities while scaling efficiently to larger models.
Congrats to the @deepseek_ai team on the launch!
BREAKING: GPT‑Image‑2.5 takes the top two spots on Image Editing Arena!
Sunburst debuts in 1st place with an Elo of 1386, followed by Flare in 2nd with 1360.
@OpenAI now holds all three top positions - and has held the top position since GPT‑Image‑2 launched in April.
GPT‑Image‑2.5 also generates edits up to 6.2× faster than its predecessor, establishing a new Pareto frontier for Speed vs. Preference.
Huge congratulations to the team!
BREAKING: Muse Spark 1.3 (xhigh) takes 1st overall on Website Arena with an Elo of 1362!
This is a jump of 5 positions from Muse Spark 1.2, establishing a new Pareto frontier for Speed and Price.
Only a month after the release of Muse Spark 1.2, @AIatMeta has topped this category on Design Arena.
Note: GPT-6 Astra is still pending final results.
Congrats to the @Meta team!
Leaderboard Updates: Muse Spark 1.3 (xhigh) by @AIatMeta takes 3rd overall on Design Arena (Elo 1367), Fullstack Arena (1315), and Web Apps Arena (1309).
With top-three performance across both non-agentic design generation and agentic web development, Muse Spark 1.3 (xhigh) is one of the strongest all-around website design models we’ve evaluated.
At an average of 187.6 seconds in non-agentic tasks, it also lands on the speed–preference Pareto frontier and establishes a new price–preference frontier.
Huge congratulations to the @AIatMeta team!
GPT-Image-2.5 Flare & Sunburst by @OpenAI are now available on Design Arena!
GPT-Image-2.5 Flare delivers higher-quality images, improved editing, and faster generation, while Sunburst is built for premium visual workflows with greater control across edits. Together, they support everything from creative content and rapid prototyping to production-ready campaign creative and polished product imagery.
Congrats to the @OpenAI team on the launch!
ChatGPT Images 2.5—faster, sharper, smarter, with better tools for creating whatever you can dream of.
- Faster image generation to keep your ideas flowing
- Improved fidelity for more natural, recognizable images
- Consistent details across multiple edits
- Comment-based edits to change only what you want
BREAKING: GPT-6 Astra (xhigh) takes #1 on Design Arena’s 3D Design leaderboard with an Elo of 1495.
It leads Kimi K3 by 65 points and improves 79 points over GPT-5.6 Sol (xhigh), OpenAI’s previous top-performing model.
A decisive new SOTA for 3D generation. Congratulations to the @OpenAI team!
GPT-6 Astra by @OpenAI is now available on Design Arena!
GPT-6 Astra brings state-of-the-art performance across computer use, browsing, software engineering, cybersecurity, science, and professional work. It’s also OpenAI’s most aligned model yet, with improved understanding of user intent and stronger judgment around the scope of delegated tasks - giving users greater confidence when handing off complex work.
Congrats to the @OpenAI team on the launch!
August was a busy month for video generation.
MiniMax H3 Max, post-trained by @fal, took the top spot on Image-to-Video, building on the strength of MiniMax H3 by @MiniMax_AI.
Wan 3.0 by @AlibabaGroup took 3rd overall, a 12-position leap from Wan 2.7.
We also saw @bfl_ai debut its first video model, FLUX 3 Video, with a strong overall performance.
These releases pushed more than just quality. H3 Max cut generation time by 46x compared to H3 in our testing, while Omni 1.1 by @GoogleDeepMind introduced cheaper, faster drafts - redefining the SOTA baseline across quality, speed, cost, and control.
Four August releases now sit in the top 11, including two of the top three.
Here’s where the Image-to-Video leaderboard stands heading into September: