🚨New LLM paper🚨
We demonstrate that automated annotation using LLMs requires validation!
Using GPT-4, we replicated 27 annotation tasks from articles in high-impact journals & show that LLM performance is promising but contingent on task/dataset (1/7)
https://t.co/kgprzT0GB0
Three exciting opportunities at
@MSFTResearch in NYC!!! 🎉
Internship w/ FATE: https://t.co/pFi2guxg8S
Postdoc w/ FATE: https://t.co/q69ur572bH
Internship w/ STAC (and collaborators in FATE) focusing on AI evaluation and measurement: https://t.co/ejuRQExgtY
Alright, people, let's be honest: GenAI systems are everywhere, and figuring out whether they're any good is a total mess. Should we use them? Where? How? Do they need a total overhaul?
Super excited to announce that @MSFTResearch's FATE group, Sociotechnical Alignment Center, and friends have several workshop papers at next week's @NeurIPSConf. A short thread about (some of) these papers below... #NeurIPS2024
Great to see everyone at LMiSS! Fantastic papers, discussions and chat. Appreciate @nick_pangakis and the UPenn crew for making this happen.
https://t.co/UAiYIVv6kz
Very little polisci work using LLMs is "replicable" in the usual sense of that term.
If this sounds like a problem, come see ppr w @cbarrie and @LexiPalmer_ Thurs @ APSA 12-130pm, PCC112A. We give some theory + empirics on where things stand and where we should be trying to go
Folks -- just an FYI that the deadline to submit for the Language Models preconference is April 19.
Also, I'm not sure why it matters, but no, we do not plan to serve La Croix at the lunch. You can BYO, but I'm not buying it for you.
cc @nick_pangakis
🚨New LLM paper🚨
We demonstrate that automated annotation using LLMs requires validation!
Using GPT-4, we replicated 27 annotation tasks from articles in high-impact journals & show that LLM performance is promising but contingent on task/dataset (1/7)
https://t.co/kgprzT0GB0
🚨 Language Model APSA Preconference🚨
Applications now open for the UPenn/Princeton LM APSA preconference (hosted by me and @nick_pangakis). Deadline is April 19. Please circulate and apply to present/discuss/attend!
https://t.co/xl36SYjw63
Just FYI for people getting APSA proposals in, me and @nick_pangakis are hosting a Large Language Models pre-conference on (day before APSA) Sept 4 at UPenn. Details will appear here, application form will open next month:
https://t.co/xl36SYjw63
An explanation for why GPT-4 is degrading:
"... we find that on datasets released before the LLM
training data creation date, LLMs perform surprisingly better than on datasets released after"
New tasks are difting away from what GPT-4 was trained on.
Excited to speak on this panel entitled, “ChatGPT turns one: How is generative AI reshaping science?” Looking forward to discussing my latest research on integrating artificial intelligence tools effectively and responsibly into research practices: https://t.co/Hh4dNj4kNV
👋NEW 📜🧵
With @ylelkes & Sam Wolken, I'm excited to release a new paper "The Rise of and Demand for Identity-Oriented Media Coverage."
Conditionally accepted @AJPS_Editor
URL: https://t.co/xyTwURK8y1
1/
A useful paper for people who want to start using LLMs in their research: AI can do a great job on data annotation, but it depends on the specific task
Especially neat is that the authors put together a tool for assessing AI annotations versus human work: https://t.co/2leKkQrr2V
To aid use case 2, we introduce a "consistency score" that empirically estimates LLM confidence in a classification. Consistency score is highly correlated w/ classification performance & can help you identify samples that may warrant closer attention from human annotators (7/7)