Our EMNLP main paper presents a fun but very challenging benchmark for LLMs: to solve over 300+ Ace Attorney detective cases that are ultra-long-context and can only be solved by long chain-of-thoughts. Here're the lessons we learned:
- Brute test time scaling does not work for long logical reasoning tasks, as GPT-4.1 can one shot the example as good as QwQ-32B
- Even when the space of possible evidences is super large, models of different sizes are still doing okay at identifying the questionable evidences
- The real struggle happens when models perform the long deduction steps, not identifying the correct facts, setting a upper bound of models' reasoning abilities
- Models tend to think a lot more for questions they are eventually wrong at, because they are stuck in a loop and just keep bringing up new points
- Surprisingly, by given the complete story context which has no additional clues, larger models do a lot better, suggesting the importance of reasoning in context
As probably zero models are benchmaxxing on deductive stories, our leaderboard faithfully shows how reasoning abilities generalize well out of domain (spoil: DeepSeek rocks).
Check it out on your model of choice!
What happens after an event boundary? We find worse memory for post-boundary items, but better source memory—a cost-benefit pattern explained by shifts in attention allocation. Our work expands Event Segmentation Theory to the start of new events.
Finally, I’m a PhD in Linguistics. Pursuing the degree in my favorite subject is a lucky and joyful journey. Great thanks to those who’ve helped me!
(can’t believe I did it by my 26).
@hxcfortuno This is amazing, DR Hong! I’m soooo happy! I’m sure you will thrive where ever you go for the next stage of your career. Any institute must be very lucky to have you!
@cogsci_soc
It’s today at language 6! Please come if you are interested in how emotion, attention, reading comprehension, and goal are intertwined together!
w @urs_maurer & Mercury & Janet, we present you…
Emotion influence beh & atten during goal-directed reading @cogsci_soc
Interested in emotional, reading comprehension, attention and how goal influence behavior? Come and check all of those in my oral!
pls rt & share the news!
@superspeeg Will definitely come and listen to your talk! I’ve seen your original post and it’s truly the most innovative approach I’ve seen in study writing system! As someone who speaks & writes in Chinese, I found this to be fascinating… looking forward!
Today we open-sourced a new project for developing behavioral experiments online. It is called Smile. Announcement of v0.1.0: https://t.co/1EsVGYMqRu... Smile has been used internally in my lab for several years and has substantially increased our productivity.