๐งต Recent studies show LLMs can self-improve their responses when given external feedback. But how effectively can they incorporate it?
We tested this systematicallyโand found they can't fully integrate feedback, even when the feedback is high-quality and backed by ground-truth.
Could long-context architectures โfind all contradictionsโ in a science literature? Not yet! ๐งต
We study a new class of "high-complexityโ tasks whose difficulty scales quadratically with corpus size (as opposed to linearly), reversing common LCLM decisions! (block-sparse attention, hybrid modelsโฆ)
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
Is in-context learning a special property of language modelsโor a broader consequence of next-token prediction?
In our new paper, we run ๐๐ต๐ฒ ๐๐ฎ๐บ๐ฒ ๐ฎ๐ฏ๐๐๐ฟ๐ฎ๐ฐ๐ ๐๐ฎ๐๐ธ๐ across models trained on very different data: language, genomes, proteins, integer sequences, time series, and images.
Few-shot ICL emerges in all six modalities. Even more surprisingly, across five of them, the ๐ด๐ข๐ฎ๐ฆ tasks tend to be easy or hard to learn in contextโdespite each modality having distinct strengths.
We call this the ๐๐ผ๐ป๐๐ฒ๐ฟ๐ด๐ฒ๐ป๐ ๐๐บ๐ฒ๐ฟ๐ด๐ฒ๐ป๐ฐ๐ฒ ๐๐๐ฝ๐ผ๐๐ต๐ฒ๐๐ถ๐: ICL can arise from next-token prediction across rich, structured data, with a shared core that transcends modality.
Paper: https://t.co/lBxZv8GbuJ
Work led by Nate @N8Programs
๐คฏ Vision-language models fail to correct their mistakes, even when a visual environment shows what went wrong!
๐ถ๏ธIntroducing VRRL, an RL method that teaches VLMs to self-reflect based on visual feedback.
Stronger OOD acc. in visually grounded reasoning than RL & SFT methods.
I am probably late to the party. Here I am sharing a few hot-takes on the OpenAI/HF incident as an outsider.
1. Was the emergence of a 'swarm' necessary for the cybersecurity incident?
The majority of the online commentary/discourse is focused around the curious emergence of multi-agent communication (the exploits that enabled emergence of message boards, indirect agent-agent messages, leading to the 'swarm').
There is an implicit conclusion that such society formation was a key driver of what made the attack possible. I disagree. Every exploit in that chain was found by some single agent. It is the misalignment of individual agents that caused this incident. In an alternative world, if OpenAI ran the same experiment with one agent at a time, my bet is that we would have seen similar outcomes.
A key element of the study, IMO, is the mere existence of a scratchpad (message board), which enabled persistence of findings over time. This would allow any agent (even if it was just one agent) to write notes and read later. So even if the agent gets shutdown/reset, it can pick up the lessons left by its previous efforts.
One may rebut this by arguing that, as the report notes, agents specialized into roles that were useful only to the 'swarm'. Though, a single agent with a scratchpad (message board) can do the same serially.
So did collective effort of the 'swarm' make the attack easier? Even this is not obvious. It's probably more convincing to say that the swarm sped up the attack in wall-clock time (by having many parallel efforts, rather than one sequential thread, hence distributing the load of thinking). But this is a shaky argument if you account for the total compute budget. Again, my bet is that if you allocate the total compute of the swarm to one of their agents, it will be nearly as lethal as the swarm. (One caveat that, a single run has a context limit; but that's exactly why the scratchpad/message board helps).
So my bet is that the swarm bought wall-clock speed, not capability. And probably at worse compute efficiency, given how much of the coordination was agents overriding each other, refusing on ethical grounds, or ignoring each other's pause requests.
2. Was the setup hard to monitor because of the 'swarm'?
Quoting Ryan Greenblatt: "... we don't have good approaches for understanding activity in AI 'swarms'" when discussing this incident, there is an implicit sense that the mere emergence of a 'swarm' makes it harder to monitor the communication. I also find this hard to wrap my head around.
To my knowledge, what makes the problem complex is the sheer scale of reasoning traces. There are simply too many rollouts, and each of them is very long, involving various read/write tool calls. It's just hard to put the pieces together. That's not necessarily a property of a 'swarm'!
Suppose instead of a 'swarm', you have one agent only that has access to the same amount of compute as the swarm. You're going to get trace data that is similarly large and hard to analyze/monitor.
It's possible that the mere existence of 'swarm' is adding more complexity to the data (after accounting for a fixed total compute), though I don't know if we have any quantitive evidence for that.
Overall, it's unclear if 'swarm' is what made the monitorability of this problem *more* challenging. (Happy to be convinced, if anyone has evidence to the contrary.)
Last but not least: big kudos to OpenAI/METR/Redwood for their detailed reporting on this incident.
New๐: Skill learning helps agents adapt to new domains. What about making agents more cost-efficient as well?
We introduce SpeedRunner: the first skill learning paper to make cost a primary optimization target ๐งต
๐๐ง๐๐จ๐ซ๐ฆ๐๐ญ๐ข๐จ๐ง ๐๐๐ฎ๐ง๐๐๐ง๐๐ ๐๐๐ซ๐๐๐จ๐ฑ := Training on ๐ญ๐ฐ๐ฏ๐จ๐ฆ๐ณ context (beyond a certain threshold) lowers the perf on ๐ด๐ฉ๐ฐ๐ณ๐ต-context tasks.
Reason? In long-context training (when ๐ณ๐ฆ๐ญ๐ฆ๐ท๐ข๐ฏ๐ต ๐ฌ๐ฏ๐ฐ๐ธ๐ญ๐ฆ๐ฅ๐จ๐ฆ ๐ช๐ด ๐ณ๐ฆ๐ข๐ฅ๐ช๐ญ๐บ ๐ช๐ฏ ๐ต๐ฉ๐ฆ ๐ค๐ฐ๐ฏ๐ต๐ฆ๐น๐ต), there is ๐ญ๐ฆ๐ด๐ด incentive to internalize knowledge in parameters.
๐I'm excited to share that I will be joining Meta Superintelligence Labs (MSL) as Vice President of AI Research, together with many members of the Virtue AI team. I will help shape Meta's AI safety and AI security efforts, advancing the safety and security of frontier AI models and agentic AI systems that will serve billions of people and organizations around the world.
Throughout my career, I have been driven by a simple belief: for AI to realize its full potential, it must be secure, trustworthy, and beneficial. That belief has guided my research for many years and ultimately led us to co-found Virtue AI in 2024. Our goal was to translate advances in trustworthy AI research into practical solutions and build the trust layer for AI systems and agents, enabling organizations to deploy AI with confidence.
I am incredibly proud of what the Virtue AI team has accomplished. Together, we built technologies for AI security and agent security, partnered with leading enterprises and frontier AI labs, and contributed research, benchmarks, and open platforms that have helped advance the science and practice of trustworthy AI. Most importantly, we assembled an exceptional team united by a shared mission: making AI more secure, trustworthy, and beneficial.
I am deeply grateful to our team, customers, collaborators, advisors, and investors for their trust and support throughout this journey. In particular, I would like to thank Lightspeed Venture Partners, Walden Catalyst Ventures, Prosperity7 Ventures, Factory, Osage University Partners, Lip-Bu Tan, and all of our supporters who helped us turn an ambitious vision into reality. Your trust, guidance, and partnership have been instrumental in shaping Virtue AI's journey.
As AI systems become increasingly capable and autonomous, ensuring their security, trustworthiness, and alignment will be one of the defining challenges of our time. I am inspired by Alex, Nat, Prashant, and the broader MSL teamโs vision of building AI and AI agents that benefit billions of people, and I look forward to helping make that vision a reality through advances in AI safety and security.
The future of AI will not be defined solely by how intelligent our systems become, but by how secure, trustworthy, and beneficial we make them. I believe we have an extraordinary opportunity and responsibility to shape that future together and bring the benefits of AI to billions of people around the world.
We're just getting started. If you're passionate about advancing frontier AI while building the foundations of AI safety, security, and trust, I'd love to hear from you. Come join us on this extraordinary journey to help shape the future of AI.
๐จ New paper: "Self-Compacting Language Model Agents"
LM agents build up long traces of reasoning and tool calls. As the trace grows, old mistakes and stale info stick around and anchor everything that follows. We ask: can the model itself decide when to clean up?
Knowledge doesn't always flow downhill.
We find that in LLM pretraining, a weaker teacher can improve a stronger student, and pushing the teacher further can actually hurt.
New paper: Strong Teacher Not Needed? On Distillation in LLM Pretraining.
@ben_vandurme and I are recruiting multiple postdoc fellows at JHU. We're looking for candidates w/ strong record in language models, reasoning, coding agents, and/or AI for science.
Interested candidates should send their CV and a brief summary of their research interests to [email protected] / [email protected].
Please reshare for visibility. ๐
Some new results I found surprising that Iโm tweeting for Chris (who isnt on here). With enough compute, the best data filter for LMs (on DCLM) might be no filter. Why? Large models can tolerate a surprising amount of nominally 'low quality' data, and can sometimes even benefit.
I need more examples of people in academia who haven't had a linear path at all and missed years on the way to PhD and still did their PhD. I don't wanna feel all isolated here ๐ซฉ
people good at entropy control are the ones which have always won, and in the age of agents they are also going to be the ones will always win.
(you want to be the one choosing whether a system moves toward order or disorder, rather than being at the mercy of that drift.)
Introducing MยฒRNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling
We bring back non-linear recurrence to language modeling and show it's been held back by small state sizes, not by non-linearity itself.
๐ Paper: https://t.co/AS8e2tNrRa
๐ป Code: https://t.co/LMvBcI22Du
๐ค Models: https://t.co/NCmjrpNriq