Congrats to Prof. @LuWang__ on receiving an Amazon Research Award from @AmazonScience for work on detecting deceptive coordination in multi-agent AI systems.
Read more: https://t.co/x6G7jLV3cs
Warm congratulations to Prof. @LuWang__ on receiving an Amazon Research Award! 🎉
This will support her project aiming to develop benchmarks and monitoring tools to detect when #AI agents coordinate around deceptive or misaligned goals. Details below:
https://t.co/qwYprsqQVi
Announcing the #AmazonResearchAwards fall 2025 recipients:
🔍 68 researchers
🏫 49 universities
🌏 11 countries
Each gains access to 800+ Amazon public datasets and AWS AI/ML tools. Meet the cohort: https://t.co/47jUdPuRrV
Alignment faking = when an AI behaves differently based on whether it thinks it's being watched 👀
Our new diagnostic (VLAF) shows this is far more widespread than previously reported.
A thread on what we found 🧵👇
https://t.co/5fFhoafmiE
@JieRuan75@LuWang__
🚀 Excited to introduce LiveOIBench
Can LLMs beat human contestants in informatics olympiads?
We evaluate LLM solutions using official test cases and compare them directly against human rankings across real Olympiad contests.
📄 Paper: https://t.co/pCqttu8Kat
🌐 Leaderboard: https://t.co/XOj67zYUdr
📊 Data: https://t.co/YLLsVkjyId
💻 Code: https://t.co/bj5XJSQtNV
🧵⬇️📷
Excited to present MLRC-Bench at #NeurIPS2025, a new suite of tasks curated from ML conference competitions for objective evaluation of AI research agents. Happy to chat!
🗓️ Thu, Dec 4 | 11 a.m. - 2 p.m. PST
📍 Exhibit Hall C,D,E #1910
🔥 Excited to introduce ManyICLBench (ACL 2025)
🧐 Do many-shot ICL tasks evaluate LCLMs' ability to retrieve the most similar examples or learn from many examples? We carefully analyzed numerous tasks and categorized them.
📄 Paper: https://t.co/9L9vTUNH2e
#ACL2025
🚨 Deadline for SCALR 2025 Workshop: Test‑time Scaling & Reasoning Models at COLM '25 @COLM_conf is approaching!🚨
https://t.co/OgWx0oElKw
🧩 Call for short papers (4 pages, non‑archival) now open on OpenReview! Submit by June 23, 2025; notifications out July 24.
Topics span RLVR, agentic test-time scaling, non-verifiable tasks, PRMs, ORMs, scaling laws, benchmarks, safety, among others.
📢 Invited speakers include:
- Aviral Kumar (CMU)
- Xuezhi Wang (DeepMind)
- Nathan Lambert (AI2)
- Lewis Tunstall (HuggingFace)
- Azalia Mirhoseini (Stanford)
Submit your papers by June 23rd.
Submission link: https://t.co/LGGF2EEgOQ
Organizers:
Lu Wang @LuWang__
Honglak Lee @honglaklee
Sewon Min @sewon__min
Sean Welleck @wellecks
Hao Peng @haopeng_nlp
Muhammad Khalifa @MKhalifaaaa
Yunxiang Zhang @YunxiangZhang4
Lifan Yuan @lifan__yuan
Shivam Agarwal @Shivamag12
See you all in Montreal!
🔍LLMs now give medical diagnoses, legal advice, and even tackle scientific problems.
❓Your LLM sounds smart. But what if it’s just good at faking expertise?
🚀We built ExpertLongBench to find out.
📉And the results? They revealed several concerns.👇
🔗 https://t.co/sfZ6UzV9Ao
🚨Just released:🚨
Proud to share Michigan AI’s white paper, proposing a comprehensive framework for how we should evaluate real-world GenAI systems, emphasizing diverse, evolving inputs and holistic, dynamic, & ongoing assessment approaches.
Now online:
https://t.co/tf6aytJLx7
🚨Announcing SCALR @ COLM 2025 — Call for Papers!🚨
The 1st Workshop on Test-Time Scaling and Reasoning Models (SCALR) is coming to @COLM_conf in Montreal this October!
This is the first workshop dedicated to this growing research area.
🌐 https://t.co/xEOVuyi473
📢New benchmark out!
We introduce CLASH, a benchmark of 345💥high-stakes dilemmas and 3,795 perspectives to evaluate how well LLMs handle complex value reasoning.
GPT-4 and Claude? Not quite there.
📄 https://t.co/EUkdH577tz
🤗 https://t.co/JlIEucJMM3
🚨 New Benchmark Drop!
Can LLMs actually do ML research? Not toy problems, not Kaggle tweaks—but real, unsolved ML conference research competitions?
We built MLRC-BENCH to find out.
Paper: https://t.co/Wdpi0SWmcq
Leaderboard: https://t.co/RMXSWGRbId
Code: https://t.co/qT9C5gJrYh
Heard of the Alaska-Hawaii merger?🤔Wonder if LLMs know it’s pending government approval before it can happen? They stumble, but we’ve got a fix⚒️!
Dive into my #EMNLP2024 work 𝐍𝐚𝐫𝐫𝐚𝐭𝐢𝐯𝐞-𝐨𝐟-𝐓𝐡𝐨𝐮𝐠𝐡𝐭—a special prompting technique to unlock LLMs’ temporal reasoning
🌍 How Verifiable Are LM Responses in the Wild? A Three-Way Factuality Benchmark
Meet 𝐅𝐚𝐜𝐭𝐁𝐞𝐧𝐜𝐡 – an updatable benchmark for evaluating language models' factuality in real-world scenarios.
🔗 https://t.co/tpjFU48m3c
@launchnlp@michigan_AI@UMichCSE
📢 NAACL needs Reviewers & Area Chairs! 📝
If you haven't received an invite for ARR Oct 2024 & want to contribute, sign up by Oct 22nd!
➡️AC form: https://t.co/4KSWkEfxoO
➡️Reviewer form: https://t.co/3DqVNOSGXF
Please RT 🔁 and help spread the word! 🗣️
#NLProc@ReviewAcl
👏A round of applause to PhD student @InderjeetNair and Prof. @LuWang__ on winning an🏆SAC Area Chair's Award - announced today at #ACL2024! Awarded to only 21 publications of the 1915 main conference papers accepted. @ACLmeeting
👉Check out their paper: https://t.co/PK3OVVLYTw
Don’t miss Frederick’s NAACL work, MOKA, exploring moral value understanding from the perspective of event-level understanding.
If AI meeting human values is your jam, this is a must-read! 🌟.
⚠️ Poster session happening soon! @FrederickXZhang@LuWang__@michigan_AI@UMichCSE
Late to join the party! Yet, stoked to share our #NAACL2024 work “MOKA: Moral Knowledge Augmentation for Moral Event Extraction”. Check out our work if interested in moral foundation theory, event extraction, news narrative, and more importantly, *aligning AI with human values*!
🏆 Heartiest congratulations to Shuyang Cao for receiving a Bloomberg Data Science Ph.D. Fellowship!
Shuyang Cao is also a member of Prof. @LuWang__'s @launchnlp Lab.
Keep up the amazing work and continue to soar high! 🎉👏
https://t.co/8XygURaQuv