We just collected over 1,300,000 human responses on SVGs and benchmarked the top 30 models. Posted it on huggingface as well, check it out! Link: https://t.co/C4vUbvrd1s
We just released the Rapidata SVG Generation Benchmark.
SVG generation is a very interesting test for frontier models: it requires visual taste, prompt following, compositional reasoning, and the ability to produce structured code that actually renders well. Additionally, we couldn't find any popular benchmark for SVGs and thought that our global human crowd would enjoy comparing SVGs instead of watching random ads.
For this benchmark, we evaluated 30 frontier LLMs on 500 static SVG prompts.
Methodology:
→ each model generated raw SVG markup
→ outputs were rasterized to 768×768 PNGs
→ humans compared model outputs head-to-head
→ results were ranked with ELO across 3 axes: Preference, Coherence, and Alignment
In total: 1,355,161 human responses.
Congrats to @joanrod_ai and the @QuiverAI team as well. Fresh off an $8.3M seed round led by a16z to build the future of vector design and visual code generation, they already rank #9 overall, alongside some of the world’s leading frontier AI labs.
If you have any benchmark you think is missing and could be useful for you, write it in comments.
Full HF dataset, including further methodology information, and weighted match results available in the comments.
#AI #Evaluation #HumanFeedback #SVG #LLMs
Rapidata SVG Benchmark just landed on ModelScope, comparing 30 frontier LLMs on static SVG generation from text prompts, with 1.35M+ human votes across preference, coherence, and prompt alignment. 🚀
🤖 https://t.co/DUNsYVKVHY
📊 Scale: 188,754 head-to-head comparisons, 500 prompts, 14,872 rasterized SVG images, and 1,355,161 human responses
🎨 Evaluation target: raw SVG markup generated by LLMs, rendered to 768x768 PNGs, then ranked by humans instead of automated metrics
🏆 Overall ranking: Claude Fable 5 Thinking leads with 1232.9 ELO, followed by Claude Fable 5 and Gemini 3.1 Pro Preview
License: CC-BY-4.0 for the benchmark prompts, with generated outputs governed by each model provider's terms.
Now that I've been through being a kid, growing up, and then having kids, it's clear that the main thing that differentiates people is simply whether they make an effort. Whether they're content to drift along with the current, or whether they try to swim.