Today, we're introducing @Intelligence_ai.
In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries.
We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities.
Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier.
People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them.
Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence.
We've helped the world's leading frontier labs break the news on their SOTA capabilities.
What's the limit? Join us and find out.
Today, we're introducing @Intelligence_ai.
In 6 months, as a team of 10, we scaled from $5M to $60M ARR and 5.5M users across 190+ countries.
We raised a $7.9M seed, led by @IndexVentures with participation from @conviction, @A_StarVC, and @combinator to build DesignArena, a universal interface for accessing and evaluating the world's AI capabilities.
Most evaluations try to simulate the real-world. We believe the real-world is the ultimate verifier.
People come to @DesignArena with a request. Models compete to fulfill the request, and users determine what works best for them.
Their live user behavior evaluates the models, improves how work is routed, and helps people access the right intelligence.
We've helped the world's leading frontier labs break the news on their SOTA capabilities.
What's the limit? Join us and find out.
BREAKING: p-image-ideogram by @PrunaAI and @ideogram_ai establishes new Pareto frontiers for both Preference vs. Price and Preference vs. Speed.
At 1K resolution, the very_low reasoning variant costs just 0.3¢ per image, while low costs 0.75¢ per image. Both generate images in around 3 seconds - roughly 8.6× faster than the Design Arena average.
A major step forward in making high-quality image generation faster and more affordable. Huge congratulations to both teams!
BREAKING: Muse Spark 1.1 takes 1st on our new Video to Website leaderboard with an Elo of 1250.
Video inputs capture richer context than static images - including interactions, transitions, and responsive behavior - challenging models to reproduce the full experience.
Only six labs currently support native video input, with @Meta debuting 14 Elo points ahead of @Kimi_Moonshot and 32 ahead of @GoogleDeepMind.
Video input is quickly becoming one of our users’ most requested capabilities.
Note: OpenAI & Anthropic currently do not yet support native video input in the API.
Congratulations to the @AIatMeta team!
BREAKING: Kimi K3 by @Kimi_Moonshot is 1st overall on 3D Design with an Elo of 1450.
This is a 6 position and 108 Elo jump from @Kimi_Moonshot's previous model, Kimi K2.6. This performance puts Kimi K2.6 82 Elo ahead of Claude Fable 5 by @AnthropicAI in 2nd and 87 Elo ahead of GLM 5.2 by @Zai_org in 3rd.
Congratulations to the @Kimi_Moonshot team on this accomplishment!
Kimi’s slides performance is quite fascinating here!
A few observations:
– Nails placements and padding (rarely cuts off elements)
– Maintains cohesiveness across the images used on slides (like looking at a Pinterest board)
– Crams a TON of information onto each slide, which can make it a bit harder to follow
– Creates great titles and storylines within individual slides, but doesn’t yet consistently nail the storyline across the full deck
The key advantage Fable 5 has at the moment is excellent focus. It keeps explanations to a minimum, understanding that the presenter should elaborate on the topic rather than letting the presentation do 100% of the work.
Still a largely unsaturated space!
BREAKING: Kimi K3 by @Kimi_Moonshot is officially 1st on Frontend Web App Arena by DesignArena
With an Elo of 1326, this open-weight model leads the way, ahead of Fable 5, Sonnet 5, and Opus 4.8 by @AnthropicAI
Huge congrats to the @Kimi_Moonshot team for this achievement!
Muse Spark 1.1 debuts with an Elo of 1189 in Python-PPTX Slides by DesignArena
The performance gap between HTML Slides highlights the capabilities differences between the two frontiers
Congrats to the team on the debut!
Muse Spark 1.1 by @Meta achieves an impressive 8-position rank increase from Muse Spark 1.0
It is the 5th lab overall on DesignArena, with an Elo of 1312
The Muse Spark model also debuts strong performance in Slides creation, data viz, and UI components - see more in the following thread
Congrats to the @Meta team for this achievement!
BREAKING: Inkling by @thinkymachines is 9th overall on Agentic Web App Arena by Design Arena with an Elo of 1257
It's an open-weight model in the same performance band as Claude Opus 4.6 by @AnthropicAI and Gemini 3.5 Flash by @GoogleDeepMind
This makes Inkling the highest-ranking US-based open-weight model for agentic workloads, achieving frontier-level performance
Congrats to the @thinkymachines team for this achievement!
Curious on how you can build your own games on Design Arena?
Learn all about our multi-file, multi-turn gamedev arena, where users can build complex single and multiplayer games, with the help of one of our founding engineers @oliverjohansson.
Today marks a historic moment for @OpenAI: GPT 5.6 Sol is officially 1st on @DesignArena.
This is the first time an @OpenAI model has held the first-place position on our single-turn HTML leaderboard.
Huge congratulations to the @OpenAI team for this achievement.
BREAKING - OFFICIAL RESULTS: GPT-5.6 Sol by @OpenAI is 1st overall on Design Arena with an Elo of 1353.
This puts GPT-5.6 Sol above Claude Fable 5 by @AnthropicAI and in the same performance band as GLM 5.2 by @Zai_org on frontend design.
This is an 18-position and 60-point Elo leap from GPT-5.5.
GPT-5.6 Sol also establishes a new Pareto frontier for preference vs. speed, faster than any model at this performance.
Congratulations to the @OpenAI team on the launch!