Solomonoff Induction looks a lot like how intelligence really works: predict by keeping every hypothesis, weighted by simplicity. Unfortunately it's impossible for any physical (finite-compute) agent to implement.
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
Our method relies on the fact that the world has universal structure, and that structure can be encoded in a program. Solomonoff induction can find that program and provably makes a bounded number of mistakes before doing so. Because it’s an uncomputable procedure, we train a neural net to approximate it without it ever seeing a single byte of natural data. It learns what it can, and through the magic of deep learning that’s enough.