Can you spot when AI bluffs?๐ค Can you outguess AIโor work with one to dominate trivia?๐
๐ We are hosting the first HumanโAI coop trivia (Quizzing) competition.
๐ฒPlay, ๐ ๏ธbuild, or โ๐ผwrite questions...
..and win prizes ๐.
๐ฅณ Itโs fun, free, and happening this June ๐ง ๐ค๐
Introducing Voice Arena STT Leaderboard.
LibriSpeech is older than the transformer.
FLEURS is people reading translated Wikipedia sentences aloud.
This is still how the industry ranks speech models. In 2026.
Today we're releasing the Voice Arena STT Leaderboard - a human-verified benchmark for STT accuracy on real, spontaneous conversation, across 7 languages: ๐บ๐ธ US English, ๐ฎ๐ณ Hindi, ๐ง๐ฉ Bangla, ๐ง๐ท Brazilian Portuguese, ๐ป๐ณ Vietnamese, ๐ท๐ด Romanian & ๐ธ๐ฆ Arabic ๐งต
US English Leaderboard โฌ๏ธ
๐ Presenting this on July 5 at 11:00 AM in Poster Session A at #acl2026nlp.
๐: https://t.co/PsnVk1Cygd
Come chat about when humans should let AI take the wheel, when they should push back, and what live human-AI collaboration reveals that static benchmarks miss.
๐งต Trusting AI is not one decision.
In our ACL 2026 paper, expert trivia players had to decide when to let AI answer, when to ignore it, and when to change their minds because of it.
The main failure mode was not blind trust.
We made a trivia game that has humans and computers cooperate: humans could let AIs buzz in for them to answer trivia questions (complete delegation) and then they worked with computers to answer other trivia questions collaboratively.
Come to #acl2026nlp Poster Session A!
๐ Huge thanks to everyone without whom this wouldn't be possible: my amazing co-authors: @YooYeonSung1, @houyu0930, @enfleisig, Irene Ying, @zhoutianyi, and @boydgraber; the UMD CLIP group; all participants and question writers/editors.
๐งต Trusting AI is not one decision.
In our ACL 2026 paper, expert trivia players had to decide when to let AI answer, when to ignore it, and when to change their minds because of it.
The main failure mode was not blind trust.
Earlier this month I successfully defended my PhD and graduated from UMD! ๐ข Thanks to everyone who played large or small role in making these last 5 years an amazing experience. Next stop, Toronto!
Defense recording: https://t.co/d24HV6HsPV
How2Everything will appear in ICML 2026! See you in Korea ๐ซก
We mine the web's procedural knowledge to better evaluate & train LLMs to generate valid step-by-step instructions, read more at:
๐ https://t.co/xzdOnwebs5
Super grateful to have four papers at #ACL2026 main! Thanks to all my co-authors from UMD, NYU, Ai2 + more ๐๐ฅน
Expect lots of "๐จ NEW PAPER" posts coming from me soon ๐
1/ ๐ย Do #LLMs really treat all languages equally when citing evidence? ๐
In our new work, we uncover linguistic nepotism: models often trade off citation quality for language preference ๐
I'll also be presenting our paper on using question-answer pairs as a new signal for spotting translation errors ๐ต๏ธ
Come to talk more about MT evaluation!
๐Poster session (Hall X4, X5)
๐Tuesday (7/29) 4-5:30pm
๐https://t.co/vRUxI3rchW