After working on LLM research and applications for about two years, I decided to go back in time and revisit some the critical works from time to time now. Takeaways in comments.
Starting with The Unreasonable Effectiveness of Recurrent Neural Networks: https://t.co/qXiZuoPF7o
1. Context is critical to forming emergent understandings in models. The whole point of RNNs is that state carries forward, letting the model remember long-range dependencies.
Our EMNLP main paper presents a fun but very challenging benchmark for LLMs: to solve over 300+ Ace Attorney detective cases that are ultra-long-context and can only be solved by long chain-of-thoughts. Here're the lessons we learned:
- Brute test time scaling does not work for long logical reasoning tasks, as GPT-4.1 can one shot the example as good as QwQ-32B
- Even when the space of possible evidences is super large, models of different sizes are still doing okay at identifying the questionable evidences
- The real struggle happens when models perform the long deduction steps, not identifying the correct facts, setting a upper bound of models' reasoning abilities
- Models tend to think a lot more for questions they are eventually wrong at, because they are stuck in a loop and just keep bringing up new points
- Surprisingly, by given the complete story context which has no additional clues, larger models do a lot better, suggesting the importance of reasoning in context
As probably zero models are benchmaxxing on deductive stories, our leaderboard faithfully shows how reasoning abilities generalize well out of domain (spoil: DeepSeek rocks).
Check it out on your model of choice!
🚀 How well can LLMs know you and personalize your response? Turns out, not so much!
Introducing the PersonaMem Benchmark --
👩🏻💻Evaluate LLM's ability to understand evolving persona from 180+ multi-session user-chatbot conversation history
🎯Latest models (GPT-4.1, GPT-4.5, o4-mini, Llama-4, Gemini 2.0, Deepseek-R1, Claude-3.7) all struggle in personalization!
🎨7 personalization skills tested in 15 scenarios
🌟Realistic long-context evaluation up to 1M tokens
👇 Check out what we discovered… (1/6)
🚀 How well can LLMs know you and personalize your response? Turns out, not so much!
Introducing the PersonaMem Benchmark --
👩🏻💻Evaluate LLM's ability to understand evolving persona from 180+ multi-session user-chatbot conversation history
🎯Latest models (GPT-4.1, GPT-4.5, o4-mini, Llama-4, Gemini 2.0, Deepseek-R1, Claude-3.7) all struggle in personalization!
🎨7 personalization skills tested in 15 scenarios
🌟Realistic long-context evaluation up to 1M tokens
👇 Check out what we discovered… (1/6)