How do LLMs really navigate the thinking space? Straight off to a final answer OR follow a wiggly path? Definitely commit OR get stuck to “infinite” self-doubting?
In our latest study, we unravel (over-)thinking through the lens of sub-thoughts: https://t.co/Wb5AIcbI6a
more in 🧵
How do LLMs really navigate the thinking space? Straight off to a final answer OR follow a wiggly path? Definitely commit OR get stuck to “infinite” self-doubting?
In our latest study, we unravel (over-)thinking through the lens of sub-thoughts: https://t.co/Wb5AIcbI6a
more in 🧵
Meet LiveOIBench‼️
THE most up-to-date and most challenging benchmark for competitive coding.
A superset of nearly every Olympiad-level dataset you’ve seen before.
🔥 Don’t miss it!
🚀 Excited to introduce LiveOIBench
Can LLMs beat human contestants in informatics olympiads?
We evaluate LLM solutions using official test cases and compare them directly against human rankings across real Olympiad contests.
📄 Paper: https://t.co/pCqttu8Kat
🌐 Leaderboard: https://t.co/XOj67zYUdr
📊 Data: https://t.co/YLLsVkjyId
💻 Code: https://t.co/bj5XJSQtNV
🧵⬇️📷
Are you interested to join @GoogleDeepMind as a student researcher (PhD students)?
I am hiring for a fun project on reward modeling for world models!
If you are interested on video understanding/generation and reward models/critiquing, get in touch and apply below! 👇
How do LLMs really navigate the thinking space? Straight off to a final answer OR follow a wiggly path? Definitely commit OR get stuck to “infinite” self-doubting?
In our latest study, we unravel (over-)thinking through the lens of sub-thoughts: https://t.co/Wb5AIcbI6a
more in 🧵
🔥 Excited to introduce ManyICLBench (ACL 2025)
🧐 Do many-shot ICL tasks evaluate LCLMs' ability to retrieve the most similar examples or learn from many examples? We carefully analyzed numerous tasks and categorized them.
📄 Paper: https://t.co/9L9vTUNH2e
#ACL2025
🔍LLMs now give medical diagnoses, legal advice, and even tackle scientific problems.
❓Your LLM sounds smart. But what if it’s just good at faking expertise?
🚀We built ExpertLongBench to find out.
📉And the results? They revealed several concerns.👇
🔗 https://t.co/sfZ6UzV9Ao
🚨Announcing SCALR @ COLM 2025 — Call for Papers!🚨
The 1st Workshop on Test-Time Scaling and Reasoning Models (SCALR) is coming to @COLM_conf in Montreal this October!
This is the first workshop dedicated to this growing research area.
🌐 https://t.co/xEOVuyi473
📢New benchmark out!
We introduce CLASH, a benchmark of 345💥high-stakes dilemmas and 3,795 perspectives to evaluate how well LLMs handle complex value reasoning.
GPT-4 and Claude? Not quite there.
📄 https://t.co/EUkdH577tz
🤗 https://t.co/JlIEucJMM3
🚨 New Benchmark Drop!
Can LLMs actually do ML research? Not toy problems, not Kaggle tweaks—but real, unsolved ML conference research competitions?
We built MLRC-BENCH to find out.
Paper: https://t.co/Wdpi0SWmcq
Leaderboard: https://t.co/RMXSWGRbId
Code: https://t.co/qT9C5gJrYh