PhD Student @UW, working on LLM research | Now @Snowflake | Prev @MSFTResearch @IntelAI @AmazonScience @UofIllinois @ZJU_China| Open for collaboration.
@thsottiaux please develop a skill so I can let my codex to run claude code for those 'simple' task in subagent without wasting my claude plan while I stay with codex : )
New GPT model is incredibly fast and so good to work with (with so much reset usage these days). If you also stuck with the slow Claude throughput, try codex : )
🚀 We built Visual Aesthetic Benchmark (VAB). Arena is alive here: https://t.co/GKSvoJ5jAk
Aesthetic judgment is one of the hardest ceilings for AI to crack right now. Not generating images, but truly understanding what “looks good.”
We hand-curated 400 sets of artist works (fine art, photography, and illustration), featuring 2000+ hours of brand-new commissioned data created specifically for this benchmark — all grounded in 13K+ domain expert judgments across 7 core aesthetic dimensions (composition, lighting, technique, expression…) to ensure rigorous evaluation in highly subjective domains.
We asked 20+ frontier AI models to judge visual aesthetics (fine art, photography, illustration) against domain experts.
Frontier models are really not good at it yet.
Best model, Claude Sonnet 4.6 hit 26.5%. Human experts: 68.9%.
> Blog: https://t.co/XvEgwkmRqr
> Leaderboard: https://t.co/BEsHCgTmjC
“Research on AI, by AI, for AI and all, shall not perish from the scientific community.”
Thrilled to share our new work BadScientist—honored with the Best Paper Award at Agents4Science 🏆.
“Research on AI, by AI, for AI and all, shall not perish from the scientific community.”
Thrilled to share our new work BadScientist—honored with the Best Paper Award at Agents4Science 🏆.
“Research on AI, by AI, for AI and all, shall not perish from the scientific community.”
Thrilled to share our new work BadScientist—honored with the Best Paper Award at Agents4Science 🏆.
In parallel with our prior work SoSBench (https://t.co/3Mo28Mnjr0) within AI4Science, we aim to contribute to more rigorous AI-accelerated science. Feedback and collaborations welcome!
🎉Thrilled to share our work got best honorable mention at ICLR BiAlign @bi_align workshop. Really appreaciate the great opportunity and thanks for the collaborative work of all authors.
#ICLR2025
1/6 🤖
Introducing SAFECHAIN—a new study on Large Reasoning Models (LRMs) with lengthy chain-of-thought (CoT). Key takeaway: more reasoning ≠ automatically safer answers. #AI#NLP#LLMSafety
While in our SafeChain paper (released on 2025/02/17), we explored similar idea first as well as LessThink and MoreThink. If you are interested in the effect of such design, don't miss it.
https://t.co/dctnLBRxd7
Qwen 3 just dropped. A very interesting design in the system is the `enable_thinking` control in the chat template.
If it is disabled, then model will be enforced to an empty thinking trace. In other words, Qwen 3 is an intrinsic reasoning model
#Qwen3#Qwen#LLM