🚀 We built Visual Aesthetic Benchmark (VAB). Arena is alive here: https://t.co/GKSvoJ5jAk
Aesthetic judgment is one of the hardest ceilings for AI to crack right now. Not generating images, but truly understanding what “looks good.”
We hand-curated 400 sets of artist works (fine art, photography, and illustration), featuring 2000+ hours of brand-new commissioned data created specifically for this benchmark — all grounded in 13K+ domain expert judgments across 7 core aesthetic dimensions (composition, lighting, technique, expression…) to ensure rigorous evaluation in highly subjective domains.
We asked 20+ frontier AI models to judge visual aesthetics (fine art, photography, illustration) against domain experts.
Frontier models are really not good at it yet.
Best model, Claude Sonnet 4.6 hit 26.5%. Human experts: 68.9%.
> Blog: https://t.co/XvEgwkmRqr
> Leaderboard: https://t.co/BEsHCgTmjC
Our new work BadScientist honored with the Best Paper Award at Agents4Science 🏆.
Many thanks to the organizers and reviewers for the recognition — and I’d love to hear your thoughts or feedback!
“Research on AI, by AI, for AI and all, shall not perish from the scientific community.”
Thrilled to share our new work BadScientist—honored with the Best Paper Award at Agents4Science 🏆.
🧩 VLMs can finally solve logic puzzles! We created VisualSphinx - a 660K synthetic vision logic dataset that boosts vision logical reasoning by 15-35% with RL. Full pipeline from 4K seed problems → 660K multimodal puzzles, all open-sourced! ✨
Website: https://t.co/aUgj1Tjcv7