We just released the Rapidata SVG Generation Benchmark.
SVG generation is a very interesting test for frontier models: it requires visual taste, prompt following, compositional reasoning, and the ability to produce structured code that actually renders well. Additionally, we couldn't find any popular benchmark for SVGs and thought that our global human crowd would enjoy comparing SVGs instead of watching random ads.
For this benchmark, we evaluated 30 frontier LLMs on 500 static SVG prompts.
Methodology:
→ each model generated raw SVG markup
→ outputs were rasterized to 768×768 PNGs
→ humans compared model outputs head-to-head
→ results were ranked with ELO across 3 axes: Preference, Coherence, and Alignment
In total: 1,355,161 human responses.
Congrats to @joanrod_ai and the @QuiverAI team as well. Fresh off an $8.3M seed round led by a16z to build the future of vector design and visual code generation, they already rank #9 overall, alongside some of the world’s leading frontier AI labs.
If you have any benchmark you think is missing and could be useful for you, write it in comments.
Full HF dataset, including further methodology information, and weighted match results available in the comments.
#AI #Evaluation #HumanFeedback #SVG #LLMs
Excited to be doubling down on @RapidataAI alongside @iaventures and @canaanpartners, after leading the pre-seed a while back. Looking forward to see them build out their real-time AI x humans RLHF cloud
Today we’re announcing an $8.5M seed round led by @canaanpartners and @iaventures , with participation from @AcequiaCapital and @blueyard.
Modern AI systems depend on large volumes of human feedback to train and evaluate — but collecting that data at scale has been slow and fragmented.
Rapidata enables teams to gather targeted feedback from real people worldwide on demand, compressing cycles that once took months into hours or days and allowing continuous model improvement.
We’ll use this funding to expand our global human data network as demand for fast, high-quality human insight continues to grow.
Turning human insight into scalable infrastructure for AI.
Announcing our $8.5M Seed
AI isn’t bottlenecked by compute anymore. It’s bottlenecked by humans.
Models still depend on large-scale human feedback to train and evaluate. Collecting that data can take weeks or months. That’s incompatible with how fast models are improving.
Spotted in the wild at ICCV 2025!
We can run such an eval actually in just a few minutes! (with 100% human responses)
GitHub page of the paper:
https://t.co/Ab3Dohx6Ll
100 founders, builders, researchers, investors and thinkers. One big question: how to solve some of the largest challenges building at the frontiers of AI. From physical compute scaling constraints to applying AI to unlock civilizational progress from material science to drug discovery. Camps, hikes and heated debates on human x AI co-existence: this was Climbing Hard AI Peaks.
Thanks to everyone for coming out and making it so special. We learned a lot from all of you. Big shout-out to our founders working on some of the biggest AI unlocks, the many researchers and builders from @cdtm_munich, @ETH_en , @Stanford and special guests from @AnthropicAI, @OpenAI , @GoogleDeepMind and more.
April 2, 2025: OpenAI releases PaperBench, testing an agents capability to replicate findings of ICML 2024 papers. Turns out Claude Sonnet 3.5 was on top back then.
August 7, 2025: OpenAI releases GPT-5 family. No PaperBench score published.
Did it not improve over o1-high?
🚀 @RunwayML's Gen-3 Alpha – The Style & Coherence King 👑
It ranks #3 overall in our text-to-video benchmark, but when it comes to style & coherence, it even beats @OpenAI Sora. 🔥
Its weakness? Alignment. ⚠️
We just dropped a new dataset on @huggingface evaluating @runwayml's Gen-3 Alpha! 📊👇
Want to benchmark your model against the biggest names? Let's talk. 💬
Aurora, the new image model from @elonmusk's @xai, just beat @openai's dalle3 model in the first ever benchmark of this model.
We collected over 401k human annotations to score this latest model that you can test on @grok.
🔍 Massive human feedback dataset for text-to-image models from @RapidataAI
- 1.5M human responses from 152K participants
- Evaluates image coherence, style & prompt alignment
- Includes detailed error heatmaps
- Covers DALL-E, Midjourney, Imagen outputs
Available on @huggingface