We just collected over 1,300,000 human responses on SVGs and benchmarked the top 30 models. Posted it on huggingface as well, check it out! Link: https://t.co/C4vUbvrd1s
Since we’ve been expanding our work to support #worldmodels, #3Dgeneration, and #robotics through live human feedback for eval and training, we decided to get a booth at #SIGGRAPH2026 in L.A. next week :).
We’d love to hear more about challenges in benchmarking and training practices for those systems.
Of course, for researchers working on the image side, we’re still very happy to chat ;)
Come say hi at booth 753!
Which shape is “Kiki”, which is "Bouba"?
Is a rubber duck considered more “reckless” than a bowling ball?
We just published a small open dataset on Hugging Face: approx. 200,000 human responses to 20 mental-association questions inspired by effects like Bouba/Kiki.
The dataset is more of a playful thing, tho it links to something a bit more central in GenAI. Many questions in AI evaluation do not have clean objective labels. They depend on perception, subjective preference, culture, and context.
We do a lot of this "subjective nature work" at Rapidata when working on evaluation or post-training for our clients. Human feedback is not only about deciding what is “correct”, but also about understanding how different people perceive outputs, and helping models align more closely with human judgement across demographics and contexts.
We just released the Rapidata SVG Generation Benchmark.
SVG generation is a very interesting test for frontier models: it requires visual taste, prompt following, compositional reasoning, and the ability to produce structured code that actually renders well. Additionally, we couldn't find any popular benchmark for SVGs and thought that our global human crowd would enjoy comparing SVGs instead of watching random ads.
For this benchmark, we evaluated 30 frontier LLMs on 500 static SVG prompts.
Methodology:
→ each model generated raw SVG markup
→ outputs were rasterized to 768×768 PNGs
→ humans compared model outputs head-to-head
→ results were ranked with ELO across 3 axes: Preference, Coherence, and Alignment
In total: 1,355,161 human responses.
Congrats to @joanrod_ai and the @QuiverAI team as well. Fresh off an $8.3M seed round led by a16z to build the future of vector design and visual code generation, they already rank #9 overall, alongside some of the world’s leading frontier AI labs.
If you have any benchmark you think is missing and could be useful for you, write it in comments.
Full HF dataset, including further methodology information, and weighted match results available in the comments.
#AI #Evaluation #HumanFeedback #SVG #LLMs
Generative models keep improving. Output diversity doesn't. Reward models are too general and ain't keeping up. #CVPR2026
Between our @RapidataAI booth, our meetup on post-training and eval, we had 400+ conversations with researchers working on image, video, and world models. A few things kept coming up:
• Almost everyone is using RL in post-training, but most approaches work with the same small set of reward models (PickScore, HPSv3, ...). These reward models are more general and mainly cover aesthetics and prompt alignment. They are useful but might not be the right objectives for many tasks.
• Model quality keeps improving, but mode collapse is on everyone's lips. Outputs are getting excellent, yet increasingly converge toward the same styles (the cat might look great, but it's too often the same cat).
• Human feedback is becoming a more central part of the training loop, with growing interest in online and continuously updated feedback rather than static datasets alone.
At Rapidata, our goal is to support this shift by making human feedback available at an adequate speed for training cycles, with 6K+ human annotations per minute.
🚨 Another thing also stood out: the field needs more benchmarks that measure areas that lab customers care about and spend the $$$ for.
We will be publishing more of these over the coming months. If there's a benchmark you'd like to see us publish (SVG generation, product position editing...), let us know in the comments.
Thanks to everyone who stopped by our booth, joined the meetup, and shared their work. Looking forward to continuing the conversations.
Bonus Pic: the robot figured I had a higher reward score
If you're attending @CVPR and curious about alternatives for typical human feedback/reward model processes in post-training/eval, drop by @RapidataAI 's booth (818). See you tomorrow! :)
Our team will be at #CVPR2026 (Booth 818). We're also hosting a meetup on GenAI post-training & evaluation. After reading lots of papers and talking to teams across labs, the same hard questions keep surfacing:
• Reward models remain imperfect proxies
• DPO/RLHF boosts quality but often hurts diversity
• Automatic video quality metrics still lag badly
• Benchmarks are saturating, eval protocols differ across labs, and good human preference data is painfully expensive
So we're hosting a meetup, next to @CVPR, for researchers and engineers to compare notes on training methodologies, evaluation design, current limitations, and where this is all heading.
Join the meetup: https://t.co/Yk5CjewqUh
-> 🤖👥 Rapidata raised $8.5M seed led by Canaan Partners & IA Ventures with Acequia Capital and BlueYard.
The platform delivers global human feedback for AI training in hours instead of weeks, turning model improvement into a continuous loop via on-demand data labeling. ⚡
-> 🧠 Led by Jason Corkill, Rapidata distributes micro-tasks across consumer apps to reach millions of users daily and match expertise to questions.
Result: real-world evaluation, faster iteration & daily model upgrades, removing one of AI’s biggest bottlenecks: human feedback. 🚀
Read More At: https://t.co/YiYk2xUcYF
Source: https://t.co/yM5e1wPKin
Today we’re announcing an $8.5M seed round led by @canaanpartners and @iaventures , with participation from @AcequiaCapital and @blueyard.
Modern AI systems depend on large volumes of human feedback to train and evaluate — but collecting that data at scale has been slow and fragmented.
Rapidata enables teams to gather targeted feedback from real people worldwide on demand, compressing cycles that once took months into hours or days and allowing continuous model improvement.
We’ll use this funding to expand our global human data network as demand for fast, high-quality human insight continues to grow.
Turning human insight into scalable infrastructure for AI.
Happiness plot twist: Sudan = happy. Poland... less so.
Global vibes are not what you think.
At least according to Rapidata's latest global happiness index. We asked 11,000 people across 110 countries how happy they are right now, with some pretty surprising results. While results like these obviously obscure some pretty important caveats (Sudan is in the middle of a brutal civil war and humanitarian crisis, and mobile phone and internet access is scarce, indicating a skew in the likely respondents, or alternatively the value of safety in conflict zones), however it does tell us something about global trends like poorer countries ranking higher (maybe money doesn't buy happiness after all).
The most interesting takeaway? The unique global reach Rapidata has and the ability to gather (near) instant global results.