@ChangHanChen3 Yes, weakly something like that. Since it's been moving in that direction for a while, that's the direction it's learning in, and all else equal you might guess it's going in that way for the future.
Transformers can learn to predict sequences of natural data without access to any data because they can be trained to approximate Solomonoff induction, the ideal next-token predictor. This is the basis of our project described in https://t.co/GYfk3RT6r5 . Here,I explain why you shouldn’t be shocked that this is possible.
A really fun project co-led with @michaelyli_ and @KfirDolev and great co-authors @gbruno_dl@ANourya@noahdgoodman, and @YoavLevine
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
Transformers can learn to predict sequences of natural data without access to any data because they can be trained to approximate Solomonoff induction, the ideal next-token predictor. This is the basis of our project described in https://t.co/GYfk3RT6r5 . Here,I explain why you shouldn’t be shocked that this is possible.
A really fun project co-led with @michaelyli_ and @KfirDolev and great co-authors @gbruno_dl@ANourya@noahdgoodman, and @YoavLevine
The result: a bias toward simplicity + learnability + tons of compute is enough. Zero-shot loss on real text, images, audio, and code falls as power laws in self-play compute, and in-context learning emerges, with the model inferring tasks from examples it's never been trained on.
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
@michaelyli_ Very cool? Does this naturally sample more diversely on high entropy tokens? In other words does this efficiently generalize the notion of branching tokens?
You're wasting FLOPs when scaling inference compute: by independently sampling parallel attempts, you burn compute rediscovering the same solutions.
Introducing QuasiMoTTo: we scale parallel sampling with correlated samples instead! These samples have higher coverage, are marginally exact draws from the LLM, and can be generated in parallel.
Result: same performance with 25-47% fewer samples in test-time scaling + 50% fewer training steps in RL!
In our new paper, we explore the design space of correlated samplers. Work with co-authors @probablynotaz9 (co-lead), @gandhikanishk, @noahdgoodman, and Emily Fox!
Can a language model learn, end-to-end, what to keep in its own KV cache and what to throw away? Can it learn to forget while it learns to reason?
Deep learning's central lesson: capability emerges from end-to-end optimization, not heuristics/strong inductive biases. But for efficiency, we rely heavily on hand-designed approaches.
🗑️ Introducing Neural Garbage Collection (NGC): we train a language model to jointly reason and manage its own KV cache, using reinforcement learning with outcome-based task reward alone. No SFT, no proxy objectives, no summarization in natural language.
New paper with @jubayer_hamid, Emily Fox, and @noahdgoodman!
Check our new AI paper, The Persian Rug!
https://t.co/C02MY9n1nW
We extract exactly the algorithm learned by the most well known model of neural network superposition, and distilled it into a set of weights resembling a Persian rug, which matches the learned loss exactly.