When questions are poorly posed, how do humans vs. models handle them? Our #ACL2025 paper explores this + introduces a framework for detecting and analyzing poorly-posed information-seeking questions!
Joint work with @boydgraber & @rachelrudinger!
π https://t.co/p4hj9K42vm
@boydgraber@rachelrudinger Check out our paper for more details!
πI'm unfortunately not at ACL this time, but our poster will be up at Poster Session 4 (Session 12) on Wednesday, July 30, 11:00-12:30, Hall 4/5!
When questions are poorly posed, how do humans vs. models handle them? Our #ACL2025 paper explores this + introduces a framework for detecting and analyzing poorly-posed information-seeking questions!
Joint work with @boydgraber & @rachelrudinger!
π https://t.co/p4hj9K42vm
@boydgraber@rachelrudinger We also discuss observations about the *process* of asking a question (e.g. asker replies are more likely when q's are poorly posed, more likely to be positive, non-dominant interpretations get better feedback, and user-refined q's drift further from dominant interpretations).
@boydgraber@rachelrudinger Both humans and models produce high-entropy interpretation distributions on poorly-posed questions - but they converge on different interpretations when questions are clear!
@boydgraber@rachelrudinger We can measure "poorly-posedness" using entropy of interpretation distributions. High entropy = answerers can't agree on what the asker wants. Low entropy = clear dominant interpretation emerges.
@boydgraber@rachelrudinger We collected 500 (question, answer, OP reply) interactions from r/NoStupidQuestions where askers gave feedback on whether their info need was met. Expert linguists annotated what interpretation each answerer actually addressed.
@boydgraber@rachelrudinger What makes a question poorly posed? When answerers can't identify a dominant interpretation despite reasoning about the asker's intent. Example: "Are chiropractors considered doctors?" β Could mean titles, training, legal powers, etc.
I'm really excited about this new line of work with my collaborators at UMD and ARL on detecting common ground misalignments (basically, misunderstandings) in human goal-oriented conversations. Great summary below, or come to @psresnik's talk today at #ACL2025NLP! 2pm Hall N.1
Linguistic theory tells us that common ground is essential to conversational success. But to what extent is it essential? Can LLMs detect when humans lose common ground in conversation?
Our ACL 2025 (Oral) paper explores these questions on real-world data.
#ACL2025NLP#ACL2025
Are you tired of using traditional stance detection to measure the polarity of text? Our #NAACL25 paper proposes an approach that uses pairwise comparisons to order texts on a continuous scale, capturing both implicit and explicit content.
πToday @ Hall 3 from 4-5:30pm.
I'll be presenting two papers @naaclmeeting!
1. Why LLMs can't write a question with the answer "468" π€π
2. A multi-agent LLM that balances opinions on "is pineapple good on pizza?" ππ
Let's also chat about:
π Helpfulness
π Why MCQA sucks
π Generating cute paper titles π
πADVSCORE won an Outstanding Paper Award at #NAACL2025@naaclmeeting!!
If you want to learn how to make your benchmark *actually* adversarial, come find me:
πPoster Session 5 - HC: Human-centered NLP
π May 1 @ 2PM
Hiring for human-focused AI dev/LLM eval? Letβs talk! πΌ
@rachelrudinger Read more at https://t.co/PFpxqyV4kv!
A huge shoutout to my advisor @rachelrudinger and everyone else in the CLIP lab at UMD for their support and feedback :-) (6/6)
I'll be presenting this work with @rachelrudinger at #NAACL2025 tomorrow (Wednesday 4/30) during Session C (Oral/Poster 2) at 2pm! π¬
Decomposing hypotheses in traditional NLI and defeasible NLI helps us measure various forms of consistency of LLMs. Come join us!
@rachelrudinger This helps us build groups of examples that evaluate the same pieces of knowledge, allowing us to measure under what *contexts* an LLM can correctly draw a particular inference ("inferential consistency"). We find that LLMs still exhibit room for improvement on this front. (5/n)
@rachelrudinger We propose a method to pinpoint the particular pieces of knowledge a defeasible reasoning example aims to evaluate by identifying the atom(s) that are most critical in determining the overall label of a defeasible NLI example. (4/n)
@rachelrudinger We also explore how atomic hypothesis decomposition can help us better understand the complexities of defeasible reasoning, a softer inference task that requires models to weigh the effects of multiple, sometimes competing, pieces of evidence on a hypothesis. (3/n)
@rachelrudinger For example, after decomposing hypothesis from an NLI premise-hypothesis pair into atoms, we can measure whether its judgment on the overall pair is consistent with its set of judgments on each premise-atom sub-problem in a logical way. (2/n)
@_joestacey_ Ah, thanks Joe!! :) And a huge thank you to you for all your early feedback -- it definitely helped the way we framed the concept of atomic inference.
1/ ποΈπ¨ Do great minds really think alike? π§ π€π¨π½βπ¦±
We investigate Human-AI Complementarity in Question Answering using CAIMIRA, our item response theory based neural framework!
w/ @haldaume3@zhoutianyi@boydgraber
Paper Link: https://t.co/1ye7fGuvdL
#nlp#human#ai#irt#qa