We are launching HALoGEN💡, a way to systematically study *when* and *why* LLMs still hallucinate.
New work w/ @shrusti_ghela* @davidjwadden@YejinChoinka 💫
🧵 [1/n]
@soldni and I are arriving to #ACL2024 🇹🇭today!
come find us at our talks & poster sessions for our OLMo & Dolma projects with @allen_ai & frens 🤩
also dont miss our poster on KIWI🥝for interactive science QA w/ our intern @brunchavecmoi & mentors @eunsolc@davidjwadden
LLMs are evaluated on the same tasks in so many different ways! 🤯
✨ We introduce OLMES – a standard for reproducible LLM evaluations that is open, practical, completely documented, and can be applied to current leaderboards & eval code bases! ✨
📜 https://t.co/SmjBV2Szsk
1/
Introducing SciRIFF, a toolkit to enhance LLM instruction-following over scientific literature. 137k expert demonstrations in 5 categories: IE, summarization, QA, entailment, and classification; models up to 70b and code to science-tune your checkpoints included! Read more in 🧵:
Looking for a dataset to enhance language model instruction-following over scientific literature? Introducing SciRIFF, a dataset of 137K expert-written demonstrations spanning 5 essential task categories for literature understanding: information extraction, summarization, question answering, claim verification, and classification. Download on HuggingFace: https://t.co/TcXRwyWTWi
Introducing our best OLMo yet. OLMo 1.7-7B outperforms LLaMa2-7B, approaching LLaMa2-13B at MMLU and GSM8k. High-quality data and staged training are key.
I am so proud of our team making such significant improvement in a short period after our first release.
Excited to share something that we've needed since the early open RLHF days: RewardBench, the first benchmark for reward models.
1. We evaluated 30+ of the currently available RMs (w/ DPO too).
2. We created new datasets covering chat, safety, code, math, etc. We learned a lot.
We hope this is a major step in understanding why reward models work, rather than just how they do for RLHF. In short, we created pairs of responses, one good one bad (with manual review) and see where reward models agree! It's a simple and powerful process.
Key takeaways:
* Running reward models is hard, we build infra to make this easier.
* We're already using this to learn more about PPO RLHF training (more on this soon).
* Reward models mirror the refusals behavior we're confused about in RLHF. Some refuse everything (including llama 2 style stuff), some refuse nothing, and few models handle both cases well.
* Datasets like Anthropic HH / Learning to Summarize only take us so far (and don't work for DPO)
* Scaling matters (big models win again)
Here's the current leaderboard:
I'm very excited about future work. Figuring out what values are reflected, generative RMs, better RMs for training, more on DPO, and everything in between.
Links!
Leaderboard: https://t.co/NZWt9tiKWt
Code: https://t.co/CmTzS0G495
Paper (arxiv soon): https://t.co/oQnXfdxdgD
Eval dataset: https://t.co/aDoxv0phD7
YouTube walkthrough: https://t.co/jWsBmXZc4P
Instruction-following capabilities of LLMs are a prerequisite to AI ✒️ writing assistance. How are good current LLMs at this task?
We present 🥝 𝗞𝗜𝗪𝗜, a dataset with instructions for knowledge-intensive, document-grounded writing for long-form answers to research questions.
📣 Job opportunities at Semantic Scholar Research @ the Allen Institute for AI (AI2) for post-doctoral & pre-doctoral researchers starting in 2024! 📣
Our team works on NLP and HCI research with a focus on open LLMs and LLM-powered research support tools and assistants.
OLMo is here! And it’s 100% open.
It’s a state-of-the-art LLM and we are releasing it with all pre-training data and code. Let’s get to work on understanding the science behind LLMs. Learn more about the framework and how to access it here:
https://t.co/utvPpWwJIp
This is fantastic news!!
Somewhat of a coincidence, but our paper that studies the effect of early arxiving on acceptance that suggested this effect is small and that it does not fill its purpose was accepted to CLeaR (Causal Learning and Reasoning) 2024
https://t.co/HVffUW7bIr
New feature alert 🚨On each paper page, scroll down to find AI-generated Topic pages related to the paper, which include topic definitions, papers most cited for the topic, and more! Now available for Computer Science fields. Here’s an example: https://t.co/OM95RBoPxz
Check out the Tulu 2 suite 🐪, a set of Llama-2 models finetuned+DPO-trained on a mixture of publicly available datasets! Our best-performing models are competitive with SoTA open models on a range of benchmarks incl. AlpacaEval and MT-Bench.
📜Paper: https://t.co/9V1VlVfogg
Using LLMs for query or document expansion in retrieval (e.g. HyDE and Doc2Query) have scores going 📈
But do these approaches work for all IR models and for different types of distribution shifts? Turns out its actually more 📉 🚨
📝 (arxiv soon): https://t.co/ObrFVZzuT1
Interested in a better way to explore #VLDB2023 papers?
Try https://t.co/soGrIXH9w8 for an LLM-powered way to probe those papers…
* Ask questions w/ a single click
* Explore answer provenance using the ending “
* Dive deep w/ recursive questions
Powered by @SemanticScholar
Does arXiving have a casual effect on acceptance?
The answer is nuanced, and depends on what assumptions you are willing to make, but arguably more importantly, we observe no difference in acceptance for different groups.
https://t.co/5SKOQkhKfY
🦙🐪🐫 So many instruction tuning datasets came out recently! How valuable are they, and how far are open models really from proprietary ones like ChatGPT?
🧐We did a systematic exploration, and built Tülu---a suite of LLaMa-tuned models up to 65B!
📜https://t.co/cFE2JUD6Zc
Today we're thrilled to announce our new undertaking to collaboratively build the best open language model in the world: AI2 OLMo.
Uniquely open, 70B parameters, coming early 2024 – join us!
https://t.co/9lQ2KYVC0v