We just released the Rapidata SVG Generation Benchmark.
SVG generation is a very interesting test for frontier models: it requires visual taste, prompt following, compositional reasoning, and the ability to produce structured code that actually renders well. Additionally, we couldn't find any popular benchmark for SVGs and thought that our global human crowd would enjoy comparing SVGs instead of watching random ads.
For this benchmark, we evaluated 30 frontier LLMs on 500 static SVG prompts.
Methodology:
→ each model generated raw SVG markup
→ outputs were rasterized to 768×768 PNGs
→ humans compared model outputs head-to-head
→ results were ranked with ELO across 3 axes: Preference, Coherence, and Alignment
In total: 1,355,161 human responses.
Congrats to @joanrod_ai and the @QuiverAI team as well. Fresh off an $8.3M seed round led by a16z to build the future of vector design and visual code generation, they already rank #9 overall, alongside some of the world’s leading frontier AI labs.
If you have any benchmark you think is missing and could be useful for you, write it in comments.
Full HF dataset, including further methodology information, and weighted match results available in the comments.
#AI #Evaluation #HumanFeedback #SVG #LLMs
Generative models keep improving. Output diversity doesn't. Reward models are too general and ain't keeping up. #CVPR2026
Between our @RapidataAI booth, our meetup on post-training and eval, we had 400+ conversations with researchers working on image, video, and world models. A few things kept coming up:
• Almost everyone is using RL in post-training, but most approaches work with the same small set of reward models (PickScore, HPSv3, ...). These reward models are more general and mainly cover aesthetics and prompt alignment. They are useful but might not be the right objectives for many tasks.
• Model quality keeps improving, but mode collapse is on everyone's lips. Outputs are getting excellent, yet increasingly converge toward the same styles (the cat might look great, but it's too often the same cat).
• Human feedback is becoming a more central part of the training loop, with growing interest in online and continuously updated feedback rather than static datasets alone.
At Rapidata, our goal is to support this shift by making human feedback available at an adequate speed for training cycles, with 6K+ human annotations per minute.
🚨 Another thing also stood out: the field needs more benchmarks that measure areas that lab customers care about and spend the $$$ for.
We will be publishing more of these over the coming months. If there's a benchmark you'd like to see us publish (SVG generation, product position editing...), let us know in the comments.
Thanks to everyone who stopped by our booth, joined the meetup, and shared their work. Looking forward to continuing the conversations.
Bonus Pic: the robot figured I had a higher reward score
If you're attending @CVPR and curious about alternatives for typical human feedback/reward model processes in post-training/eval, drop by @RapidataAI 's booth (818). See you tomorrow! :)
Our team will be at #CVPR2026 (Booth 818). We're also hosting a meetup on GenAI post-training & evaluation. After reading lots of papers and talking to teams across labs, the same hard questions keep surfacing:
• Reward models remain imperfect proxies
• DPO/RLHF boosts quality but often hurts diversity
• Automatic video quality metrics still lag badly
• Benchmarks are saturating, eval protocols differ across labs, and good human preference data is painfully expensive
So we're hosting a meetup, next to @CVPR, for researchers and engineers to compare notes on training methodologies, evaluation design, current limitations, and where this is all heading.
Join the meetup: https://t.co/Yk5CjewqUh
Join @RapidataAI at #CVPR2026 (Booth 818) we are also hosting a meetup on GenAI post-training & evaluation.
After reading lots of papers and talking to teams across labs, the same hard questions keep surfacing:
• Reward models remain imperfect proxies
• DPO/RLHF boosts quality but often hurts diversity
• Automatic video quality metrics still lag badly
• Benchmarks are saturating, eval protocols differ across labs, and good human preference data is painfully expensive
So we're hosting a meetup, next to @CVPR , for researchers and engineers to compare notes on training methodologies, evaluation design, current limitations, and where this is all heading.
Join the meetup: https://t.co/a3WLXqc0lx
Today we’re announcing an $8.5M seed round led by @canaanpartners and @iaventures , with participation from @AcequiaCapital and @blueyard.
Modern AI systems depend on large volumes of human feedback to train and evaluate — but collecting that data at scale has been slow and fragmented.
Rapidata enables teams to gather targeted feedback from real people worldwide on demand, compressing cycles that once took months into hours or days and allowing continuous model improvement.
We’ll use this funding to expand our global human data network as demand for fast, high-quality human insight continues to grow.
Turning human insight into scalable infrastructure for AI.