@emnlpmeeting@jgwinnup Hi, this blog post is still a bit confusing and overly harsh sounding. Is it meant to read like this? "if ANY author has at least 3 *CL pubs, then AT LEAST ONE author must review AT LEAST 3 papers, or else desk-reject; exception applies if NO authors have AT LEAST 3 papers"
It was fun watching GPT-4 fail miserably on this dataset! A truly herculean effort by @idavidrein and co means the questions are super high quality and interesting. I hope to see many scalable oversight evaluations on this soon...
๐จ New paper alert! ๐จ
We all know that LLMs can be integrated with tools and they can be personalised through clever prompts or more recently "custom instructions"?
But what happens when we evaluate these instructions more systematically?
https://t.co/AyrT3mW3h9
(1/N)
We are so grateful to win the Best Paper Award at #AACL2023 ๐ฅณ๐
Our work โSmoothing Entailment Graphs with Language Modelsโ bridges symbolic and explainable NLI into the era of language models.
Thank you to my coauthors, @EdinburghNLP , and @aaclmeeting !
๐ Best Paper! ๐
The 'Smoothing Entailment Graphs with Language Models' paper by Nick McKenna, Tianyi Li, Mark Johnson, and Mark Steedman wins the Best Paper Award at #AACL2023!
๐ Big congrats! ๐
https://t.co/W86IeexBuM
๐ Best Paper! ๐
The 'Smoothing Entailment Graphs with Language Models' paper by Nick McKenna, Tianyi Li, Mark Johnson, and Mark Steedman wins the Best Paper Award at #AACL2023!
๐ Big congrats! ๐
https://t.co/W86IeexBuM
These models fail to generalize in ways a human would. We offer our experimental controls as tools to evaluate future models for their ability to generalize in NLI tasks beyond artifacts seen in training.
Better late than never:
We (@TianyiLi_Teddy7 and I) are excited to announce our paper will appear in #EMNLP23 findings: "Sources of Hallucination by Large Language Models on Inference Tasks"
Preprint link: https://t.co/AbyjjNgqqj
@emnlpmeeting@InfAtEd@EdinburghNLP
LLM performance appears very good when sample labels conform to these biases, but performance drops drastically when samples do not conform. We show that these models use such biases to approximate language inference, but these biases have no relation to language meaning.
Very pleased to announce that our paper, "Smoothing Entailment Graphs with Language Models" has been accepted to #AACL2023 ๐
Many thanks to co-authors @TianyiLi_Teddy7, Mark Johnson, and Mark Steedman
Preprint link: https://t.co/ZkeBeCPRcn
Our paper explores this question with a symbolic model for NLI, an Entailment Graph.
An off-the-shelf language model is used to approximate statements missing from the EG using known statements, completing a transitive chain. This improves both recall, and even precision.