I see, so it encourages exploitation. Have you thought about how to scale exploration with compute (you already tried ensemble)? For example, the model might learn shallow circuits (eg memorization) early on which then gets reinforced by the reward. It is unclear to me if the KL to Solomonoff prior (which is the only bias toward generalization in your algorithm) is sufficient to cure this...
@AdityaCowsik Might be a stupid question. How does the alignment between gradient and recent weight change induce the learnability bias? It works for trivial cases, eg random train data, but when SNR is non-zero, do you have in mind some linear direction in weight space is the most learnable?
Can an LM, starting from random init (!!), learn to generate all of its pretraining data?
Introducing Self-Play Pretraining with Zero Data. Two models start from random initialization: a generator proposes programs for a universal Turing machine and a learner trains on their outputs. We never train on any real data, but see predictable scaling on natural datasets: zero-shot val loss on images, text, audio, and melodies decreases predictably with self-play compute. And the learner develops in-context learning capabilities.
A fun proof-of-concept, co-led with @AdityaCowsik and @KfirDolev and co-authors @gbruno_dl, @ANourya@noahdgoodman, and @YoavLevine.
@jayden_teoh_ The argument for why action chunking (AC) is important for robotics relies on “short-sightedness”, right? The same argument won’t work for LLMs because they already have the whole history in context. If MTP ends up being essential for LLM, the insights must be different I think?
@ar0cket1 Okay you’re right. I was thinking too naively that lr should inversely scale as variance when noise dominates. But since SNR is the same ie mean also increases, that was wrong. Still, my instinct is that hard prompts’ gradients are noisier so you would need some tricks there…
@akarshkumar0101 Nice work! Is it correct to say that SMT throws away transformer and keeps transition model at test time, while NextLat does the opposite? Both have morally the same training objectives.
How far can we compress billion-parameter LLMs? We introduce requential coding, which achieves < 1-bit per param compression, and explains why scaling doesn't hit a generalization wall!
https://t.co/gUZekHiFRU
w/@m_finzi, @YujiaZheng9 ,@kunkzhang, @andrewgwils
1/🧵
@DmitryRybin1 If you are willing to make the assumption K(f)~N^a, why not just assume the same for K(D|f) and then you get scaling laws for free? It feels like any reasonable argument about why K complexity being power-law would be the actual juice.
@rosinality@guilhermeotina No, the paper you cited didn’t tune weight decay. As shown in https://t.co/hrToeVjqTT, simply tuning wd improves data efficiency by 3x. If you compare numbers between papers, AR is as data efficient as diffusion. Put simply, the paper you cited has an undertuned AR baseline.
In modded-nanogpt we also found that the last couple attention layers hate interacting with the final prediction MLPs. So we work around it with a cached activation from earlier. In the attention residuals paper, Kimi doesn't explicitly mention it, but you can see from their chart that the final attn layers dont engage with the final outputs. So I think there is something fundamental going on here.
I lost track of time again >.< I'm really sorry if you DMed me lately. I promise to go over my DMs!
---
This sprint, I built a Lean4-to-TileLang Tensor Program Superoptimizer. With this, I now have a formal infrastructure where I (or my agents) can define neural network architectures in Lean4, and automatically get:
1. Optimized IO-aware accelerator kernels in TileLang. It can find FlashAttention2, FlashNorm, split-k matmul, and others automatically. I'm currently getting a ~1.8x geomean speedup on my benchmark set on A100s.
2. Optimizer choices and parametrization that enable hyperparameter transfer across width and depth (see my previous blog posts).
3. Hyperparameter scaling laws that tell us how to adjust hyperparameters as we scale batch size, training horizon, dataset size, and etc. (see quoted tweet).
4. Low-rank proxies for the optimizers to speed up hyperparameter tuning at small scales and have them transfer to the full-rank case (we have an upcoming paper on this, stay tuned!).
@LIGO@ego_virgo@KAGRA_PR Event GW250114, with signal-to-noise ratio of 80, shows
• mass & spin of the merged BH matches the Kerr spectrum
• horizon area of the final BH is greater than the sum of the horizon areas for the merging BHs, as predicted in 1971 by Stephen Hawking https://t.co/iINXwIDyq4
@Machina_Ratio@sea_snell That’s not true. GDM had several interesting papers on this, for example. Also this tutorial starting around 29mins, https://t.co/HLEMwSAYat
"A calculator app? Anyone could make that."
Not true.
A calculator should show you the result of the mathematical expression you entered. That's much, much harder than it sounds.
What I'm about to tell you is the greatest calculator app development story ever told.
We are excited to release a short course on AGI safety! The course offers a concise and accessible introduction to AI alignment problems and our technical & governance approaches, consisting of short recorded talks and exercises (75 minutes total). https://t.co/8AhEBb3Vpk