The AllenNLP team works on language-centered AI that equitably serves humanity. We deliver high-impact research and open-source tools to accelerate progress.
Olmo 3 afterglow - want to share how I came to join Ai2 to encourage others interested in LLM research. I found it difficult to break into the field as someone who does not hold a phd. Research roles at top labs are highly competitive and I didn’t have professional experience training llms or the right connections. I spent a lot of time reading research and contributing to oss projects. Ended up cold emailing @natolambert who got me an interview!
Incredibly proud of the OLMo team! Alongside the new model releases, there’s a wealth of material for the community: a full research report detailing the entire training process, the Dolma 3 dataset, all intermediate checkpoints, and the complete training scripts.
Today's the day.
Olmo 3 is out in the open 🐮 🦖
I've had a wonderful time working on this model.
I focused on post-training Olmo 3 to write code. I wrote up some of my thoughts on my blog here:
https://t.co/2m7ZvMzxDc
This release has SO MUCH
• New pretrain corpus, new midtrain data, 380B+ long context tokens
• 7B & 32B, Base, Instruct, Think, RL Zero
• Close to Qwen 3 performance, but fully open!!
OpenAI's blog (https://t.co/VeNI85798G) points out that today’s language models hallucinate because training and evaluation reward guessing instead of admitting uncertainty. This raises a natural question: can we reduce hallucination without hurting utility?🤔
On-policy RL with our Binary Retrieval-Augmented Reward (RAR) can improve factuality (40% reduction in hallucination) while preserving model utility (win rate and accuracy) of fully trained, capable LMs like Qwen3-8B.
[1/n]
We're starting to hire for our 2026 Olmo interns! Looking for excellent students to do research to help build our best models (primarily enrolled in Ph.D. with experience or interest in any area of the language modeling pipeline).
OlmoEarth is a great way to show how Ai2 investing heavily in core modeling capabilities can have positive second order effects in scientific domains.
It is a multimodal, spatio-temporal model built on a fork from the same pretraining codebase with use for text olmos.
💡Beyond math/code, instruction following with verifiable constraints is suitable to be learned with RLVR.
But the set of constraints and verifier functions is limited and most models overfit on IFEval.
We introduce IFBench to measure model generalization to unseen constraints.
Introducing IFBench, a benchmark to measure how well AI models follow new, challenging, and diverse verifiable instructions. Top models like Gemini 2.5 Pro or Claude 4 Sonnet are only able to score up to 50%, presenting an open frontier for post-training. 🧵
📢 Can LLMs really reason outside the box in math? Or are they just remixing familiar strategies?
Remember DeepSeek R1, o1 have impressed us on Olympiad-level math but also they were failing at simple arithmetic 😬
We built a benchmark to find out → OMEGA Ω 📐
💥 We found that although very powerful, RL struggles to compose skills and to innovate new strategies that were not seen during training. 👇
work w. @UCBerkeley@allen_ai
A thread on what we learned 🧵
New updates for olmOCR, our fully open toolkit for transforming documents (PDFs & images) into clean markdown. We released:
1️⃣ New benchmark for fair comparison of OCR engines and APIs
2️⃣ Improved inference that is faster and cheaper to run
3️⃣ Docker image for easy deployment
We enabled OLMoTrace for Tülu 3 models! 🤠
Matched spans are shorter than for OLMo models, bc we can only search in Tülu's post-training data (base model is Llama). Yet we thought it'd still bring some value.
Try yourself on the Ai2 playground -- https://t.co/xGDdIR99De
As we’ve been working towards training a new version of OLMo, we wanted to improve our methods for measuring the Critical Batch Size (CBS) of a training run, to unlock greater efficiency, but we found gaps between the methods in the literature and our practical needs for training OLMo. 🧵
Super excited that our second reward model evaluation is out. It's substantially harder, much cleaner, and well correlated with downstream PPO/BoN sampling.
Happy hillclimbing!
Huge congrats to @saumyamalik44 who lead the project with a total commitment to excellence.