🍰 another slice of CAKE just dropped: blackwell KDA kernels, ~2.05x geomean speedup vs FlashKDA:
https://t.co/hQBCTLUmXR by @avyyhuang
more cake kernels are baking 👀 not saying what CAKE actually is yet though, the full reveal is in the oven too.
🔥
𝐍𝐞𝐰 𝐛𝐥𝐨𝐠: 𝐓𝐨𝐰𝐚𝐫𝐝𝐬 𝐋𝐨𝐨𝐩𝐞𝐝 𝐌𝐨𝐝𝐞𝐥𝐬 𝐃𝐨𝐧𝐞 𝐑𝐢𝐠𝐡𝐭 — 𝐏𝐚𝐫𝐭 𝐈
Looped models reuse the same weights across depth, promising a better compute–parameter trade-off, especially for reasoning.
𝐁𝐮𝐭 𝟏) 𝐝𝐨 𝐭𝐡𝐞 𝐠𝐚𝐢𝐧𝐬 𝐬𝐮𝐫𝐯𝐢𝐯𝐞 𝐰𝐡𝐞𝐧 𝐛𝐨𝐭𝐡 𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐚𝐧𝐝 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝐅𝐋𝐎𝐏𝐬 𝐚𝐫𝐞 𝐦𝐚𝐭𝐜𝐡𝐞𝐝? 𝟐) 𝐀𝐧𝐝 𝐰𝐡𝐢𝐜𝐡 𝐚𝐫𝐜𝐡𝐢𝐭𝐞𝐜𝐭𝐮𝐫𝐚𝐥 𝐜𝐡𝐨𝐢𝐜𝐞𝐬 𝐚𝐜𝐭𝐮𝐚𝐥𝐥𝐲 𝐦𝐚𝐭𝐭𝐞𝐫?
We run 𝐚𝐩𝐩𝐥𝐞𝐬-𝐭𝐨-𝐚𝐩𝐩𝐥𝐞𝐬 ablations spanning Ouro to Huginn. Huginn performs better overall, with the largest gains coming from the loop-in-the-middle (sandwich) design and input injection, though they provide different benefits.
Trained on 𝟓𝟎𝟎𝐁 tokens, an 𝟖𝐁-𝐀𝟎.𝟖𝐁 Huginn MoE approaches or surpasses a 𝟑𝟐𝐁-𝐀𝟑.𝟐𝐁 feedforward MoE on several reasoning benchmarks, including GSM8K (83.6% vs. 80.8%), while using 𝟕𝟓% 𝐟𝐞𝐰𝐞𝐫 resident parameters under 𝐦𝐚𝐭𝐜𝐡𝐞𝐝 𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠 𝐚𝐧𝐝 𝐢𝐧��𝐞𝐫𝐞𝐧𝐜𝐞 FLOPs.
More details and the blog link in the thread ↓
🚀 OpenRSI is a new open research series from @FrontisAI for concrete, testable progress toward recursive self-improvement (RSI).
As its first project—and also my first work as first author—I’m proud to present OpenMLE: an open full-stack AI4AI system for autoresearch, where evolutionary agents improve ML solutions through executable feedback.
OpenMLE has three components:
- OpenMLE-Gym: 5,758 executable tasks + evaluators
- OpenMLE-ERL: execution-grounded SFT + RL
- OpenMLE-Evo: experience-guided long-horizon search
🏆 The full system—our trained Frontis-MA1-35B model paired with OpenMLE-Evo-Max—reaches 71.21% Medal Average on MLE-Bench Lite: surpassing GPT-5.5 + Codex (68.18%) and just 1.52% from GPT-5.6 Sol + Codex and the 2.8T Kimi K3 + Claude Code (72.73%). Budget: 12 hours/task on one RTX 4090 capped at 12 GB VRAM.
🌍 On 10 held-out NatureBench Lite tasks, both components transfer:
• same framework, model swap: Match-SOTA 50% → 70%
• same base model, framework swap: Match-SOTA 20% → 50%
🔓 Paper, code, models, data, and analysis below. 🧵
C2KV combines KV cache compression and non-prefix reuse for long-context LLMs, achieving up to 17x inference speedup while preserving quality.
https://t.co/KbDlqK5903
#MachineLearning#AI#LLM#DeepLearning#AgenticAI
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here.
The basic issue is that the labs are highly unprofitable. This may seem like a simple point, but private market valuations can be relatively irrational; however, like with $SPCX, post-IPO pricing will likely be much more punishing, especially as the standard 6-month lockup period expires and selling pressure intensifies.
Many people claim that the labs have high margins. Yet even with high margins, a valuation of $1T would be justified only if the labs were doing nothing aside from serving inference (thus reducing costs only to those relevant to inference) and posting annual revenue numbers in the $100-200 billion range assuming ~80% gross margin and a 20x earnings multiple.
This assumption is obviously not true, because the frontier labs have to continually spend money training the next generation of models. This is because of market competition from runner-up firms. For example, if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training, in a way such that at any given point of time, the amount you're forced to invest in the next model is dramatically higher than the amount of money you're actually making, because even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners.
(There is also a related subpoint where frontier labs claim they can distill their leading models to win out at lower intelligence levels as well. This makes no sense because the revenue numbers involved are far too low when taking into consideration the rather low margin of such inference.)
Frontier lab valuations appear largely to be based on the assumption that as you scale up, the capabilities which emerge will be sufficiently general and profound that we'll see explosive growth (https://t.co/RqmkltVpM3) from things akin to AI agents starting and autonomously managing entire companies of subagents. But it's not clear to me that this is the case; indeed, as I mentioned in my previous post (https://t.co/3URAcJ4XkJ), I believe that capabilities growth will be slower, spikier, and more data-limited than people currently assume. It may be the case that eventually we will see explosive growth of this nature with full automation of the economy, but at the very least my viewpoint implies much longer (multi-decade) timelines until we reach this point. It is not clear to me that the frontier labs will be able to operate unprofitably for so long, although I suppose maybe this foreshadows some sort of inevitable nationalization.
I also want to make a broader point about technological diffusion. The reason why technological diffusion is slow isn't just because, e.g., old people take a long time to learn how to use technology (although this is of course a contributing factor to some degree). In my view, it's because when a new, revolutionary technology comes along, the ways to incorporate that technology into subsequent developments are not always obvious, and in fact they cannot necessarily be arrived at through the application of pure reason. If they could be, then perhaps frontier models, at a certain point, would have a perfect understanding of how the LLM application layer should be developed, and they would then autonomously code, deploy, and sell such a layer.
But it seems more plausible to me that this diffusion is limited moreso by the hard problem of economic calculation--that is to say, the Hayekian notion through which the price system gradually promotes efficient allocation of resources and which cannot be simulated through central planning--and that even if we froze current capability levels at today's levels, it would take well over two decades to fully integrate in LLMs into our lives. Such a view is consequently rather bearish for the continued profitability of labs as it reduces their prospects for finding, say, something else comparable in profitability to coding agents, which seems to have been a somewhat lucky discovery by Anthropic to begin with. That is to say, even if you spam FDEs you aren't necessarily going to be able to just figure out the "correct" product shapes fast enough.
Overall, I don't think that people have clearly reasoned through their mental models for why lab equity should be worth as much as it currently is, and that if you actually bother to write down such a model, you may not arrive at the conclusion that you want to arrive at. This isn't to say that I don't expect AI to experience a huge (industry-wide) boom in the coming decades, but just that I'm not entirely sure I would buy OpenAI or Anthropic stock at latest valuations if I were given the opportunity to do so.
Of course, as an ex-lab employee, arguably this is talking against my own book; I should really be giving people more reasons to be bullish. But in the end, my influence is so small that it doesn't make a difference, so why not have some fun?
This episode dives systematically into the co-design of models and infrastructure.
Kaichao You shared a vivid analogy: comparing tokens to electricity.
Hardware represents natural resources such as wind and water power.
Models are power generators, including wind turbines and hydroelectric generators.
The inference engine acts as the power grid system.
The depth of co-design dictates power generation efficiency. For areas with rapid water currents, what generator design can best harness this energy?
You Kaichao: vLLM, Open-Source Infra, Model Co-Design & Journey from Com... https://t.co/V5SYBtNeYj 来自 @YouTube
Introducing Online KL Shampoo (OKLS), an optimizer that brings a KL-optimal approximation of full-matrix AdaGrad to language-model training.
Diagonal optimizers ignore correlations between gradient coordinates. Full-matrix AdaGrad captures this geometry but requires quadratic state. Muon considers correlations but not their history. OKLS closes this gap using KL-optimal Kronecker factors, whitening matrix gradients across both row and column directions while remaining naturally scale-invariant.
The main challenge is computing fresh inverse-square-root preconditioners at every step. Even one-step staleness can destabilize training. We make zero-staleness preconditioning practical with Scaled CANS Coupled Newton–Schulz: 10 iterations, 27 FP16 GEMMs, and FP32 accumulation.
OKLS achieves 1.45× the parameter efficiency of Muon while retaining 98% of its training throughput. Across 200M–1B models, an OKLS model matches a Muon model roughly 1.5× larger.
Lecture 10 of my course! Nominally on regularization in RL, so I discuss the evolving role of the KL penalty in RL, but also a set of nice RL papers that explain what RL helps models generalize better than SFT -- with theory supporting it.
When going through these, it's so interesting how seasonal problems in ML are. Lots of problems from controlling reward models overopt will rhyme as we try to control rubrics for agents.
00:00 Intro & the role of regularization
02:50 The KL penalty in RL
10:53 RL as a reverse KL loss
21:09 Why RL generalizes better than SFT
25:15 Other regularization tools
Just a few videos left as I get to the end of the course. Thanks all, and keep sending questions. Spread the word if you have a second.