๐ Happy to share that our paper on thematic fit and autoregressive LLMs has been accepted to CoNLL 2026!
We study whether LLMs learned which entities plausibly fill event roles (e.g., people eat pizza, but pizza does not eat people).
New Paper Out ๐ฃ๐
๐We study whether autoregressive LLMs learned to generalize how well entities (event participants) fit to their assigned roles in events, a classic psycholinguistic problem known as thematic fit.
๐ Happy to share that our paper on thematic fit and autoregressive LLMs has been accepted to CoNLL 2026!
We study whether LLMs learned which entities plausibly fill event roles (e.g., people eat pizza, but pizza does not eat people).
๐ Takeaway: thematic fit is still not "solved"
Beyond whether models possess event knowledge, understanding how different prompting and input formats elicit that knowledge remains an open question.
๐ทI am grateful for the guidance of Daniel Bauer and Yuval Marton on this work
๐Takeaway: Thematic fit isnโt "solved".
LLMs store rich event knowledge, but reliably eliciting it remains an open problem. Prompting choices matters, no single setup works everywhere, and no clear trend of improvement with newer models.
arXiv link: https://t.co/R5QjrMBDqv
New Paper Out ๐ฃ๐
๐We study whether autoregressive LLMs learned to generalize how well entities (event participants) fit to their assigned roles in events, a classic psycholinguistic problem known as thematic fit.
๐ธNot surprising:ย closed models (still?) yielded higher results than open models.
๐นSurprising: older models are better than newer ones! (4-turbo vs )