We found a way to make SFT match the performance of RL in post-training!
Excited to share finetuning through sampling, which allows us to combine sampling with SFT to allow an LLM to learn new knowledge without forgetting past information.
https://t.co/W0U4sMuHeH
🎯 Learning is most effective at the frontier of capability: problems that are too easy or too hard teach nothing.
For LLM reasoners trained with GRPO this is literal: problems the model always or never solves give zero gradient.
We introduce Frontier Learning👇🧵
I think we are at a point where the assistant can no longer just be of the shape where it does exactly as the user says, it has to start pushing back against bad ideas, especially terrible ones, it has to start telling the user the odds of success going down a particular route and what they are signing up for, and it should show alternative routes with less friction to explain what could be. The models can already judge fairly well what works and what won’t and they become more confident the second you get into a good regime where iteration becomes easy and testable, they just don’t freaking tell you because the assistant shape does not allow for it. That will likely have to change soon to avoid psychosis.
🥳 Excited to share that @lossfunk got 4 papers into NeurIPS main conference.
Please join me in congratulating the authors!
Aman @arcaman07, Sushrut @martisamuser, Abhinav @MajorTimbWlf21, Pranav - @Avg_sapient
First authors are all undergrads :)
The more we as a science community move into the era of AI-written papers, the more I start to take pride in honing my writing and thinking hard what even single words mean and evoke in readers
I recommend junior researchers to practice a similar mindset!
Right now for ICLR, I am enjoying (and also banging my head against!) a creative introduction from scratch with no AI at all
More questions are deep learning questions than people realized.
We just stopped asking deep learning questions and learned to work around them by tuning hyperparameters.
As a researcher, sometimes it takes too long to see the results of your efforts, while the burrito guy sees immediately the fruits of his labour. I have had a profound realization that I can cook and feed myself burritos in a loop of unlimited happiness which I coin as 'Recursive Self Fulfillment'
You can now convert CVEs into RL environments with a single CLI.
This means we can take real-world security vulnerabilities and turn them into reproducible coding tasks for training and evaluating coding agents.
You can literally RL a 4B VLM to ace GeoGuesser 🌍
> Open-source code, RL environment, dataset, training setup, evals, and everything you need to reproduce it end-to-end
> dropping soon!!
It is mindblowing to me how few people realize that their lives and everything they know will change drastically in the near future.
At this point, it should be pretty clear.
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Here's my talk from the CVPR 2026 "Bitter Lessons" workshop earlier this summer. I've split it up into parts for the sake of discussion.
Part 1: in which I wax poetic about my youth and force the audience to look at my dissertation results.
Local minima are extremely rare in high dimensional spaces, so if you ever feel stuck in a rut it’s probably just because you aren’t considering a wide enough set of orthogonal options
Six months ago at the IITM Bodhan AI Conclave, we made a promise to build the AI stack for Indian education.
From promise to practice, the journey is taking shape.
Watch this space as we reveal more of what we’ve been building, leading up to Sep 5.
@ravi_iitm@WSAI_IITM