Principal AI Architect @ASMLcompany. Building Foundation Models and Agentic AI in the semiconductor industry. Previously PhD in ML/Optimization @TUDelft
Three days ago I left autoresearch tuning nanochat for ~2 days on depth=12 model. It found ~20 changes that improved the validation loss. I tested these changes yesterday and all of them were additive and transferred to larger (depth=24) models. Stacking up all of these changes, today I measured that the leaderboard's "Time to GPT-2" drops from 2.02 hours to 1.80 hours (~11% improvement), this will be the new leaderboard entry. So yes, these are real improvements and they make an actual difference. I am mildly surprised that my very first naive attempt already worked this well on top of what I thought was already a fairly manually well-tuned project.
This is a first for me because I am very used to doing the iterative optimization of neural network training manually. You come up with ideas, you implement them, you check if they work (better validation loss), you come up with new ideas based on that, you read some papers for inspiration, etc etc. This is the bread and butter of what I do daily for 2 decades. Seeing the agent do this entire workflow end-to-end and all by itself as it worked through approx. 700 changes autonomously is wild. It really looked at the sequence of results of experiments and used that to plan the next ones. It's not novel, ground-breaking "research" (yet), but all the adjustments are "real", I didn't find them manually previously, and they stack up and actually improved nanochat. Among the bigger things e.g.:
- It noticed an oversight that my parameterless QKnorm didn't have a scaler multiplier attached, so my attention was too diffuse. The agent found multipliers to sharpen it, pointing to future work.
- It found that the Value Embeddings really like regularization and I wasn't applying any (oops).
- It found that my banded attention was too conservative (i forgot to tune it).
- It found that AdamW betas were all messed up.
- It tuned the weight decay schedule.
- It tuned the network initialization.
This is on top of all the tuning I've already done over a good amount of time. The exact commit is here, from this "round 1" of autoresearch. I am going to kick off "round 2", and in parallel I am looking at how multiple agents can collaborate to unlock parallelism.
https://t.co/WAz8aIztKT
All LLM frontier labs will do this. It's the final boss battle. It's a lot more complex at scale of course - you don't just have a single train. py file to tune. But doing it is "just engineering" and it's going to work. You spin up a swarm of agents, you have them collaborate to tune smaller models, you promote the most promising ideas to increasingly larger scales, and humans (optionally) contribute on the edges.
And more generally, *any* metric you care about that is reasonably efficient to evaluate (or that has more efficient proxy metrics such as training a smaller network) can be autoresearched by an agent swarm. It's worth thinking about whether your problem falls into this bucket too.
Adversarial Poetry as a Universal Single-Turn
Jailbreak Mechanism in Large Language Models
"Adversarial poetry functions as a...
jailbreak technique for LLMs. Converting ...harmful prompts into verse produced attack-success rates up to 18 times higher than their prose baselines"
metaTextGrad: Automatically optimizing language
model optimizers
"we propose metaTextGrad, which focuses on designing a meta-optimizer...our approach consists of two key components: a meta prompt optimizer and a meta
structure optimizer...improvement of up to 6% [over] baseline"
Diffusion Transformers with Representation Autoencoders
“In this work, we explore replacing the VAE with pretrained representation encoders (e.g., DINO, SigLIP, MAE) paired with trained decoders, forming what we term Representation Autoencoders (RAEs).”
@simonw Hilarious! You could try flagging this by embedding prompts and catching off topic content via distance (e.g. between system vs user prompt). You still won’t know which is the real instruction vs. the adversarial one, but at least you’d spot something odd
@rasbt@bgurley Odd that he argues users don’t care about where the model is hosted and where the inference takes place, that’s literally one of the first things enterprises need to worry about.
@arthurmensch@ASMLcompany Congrats! Looking forward to working together and applying cutting edge AI to some of the world’s hardest engineering problems!
We just announced a strategic partnership with Mistral AI, with the goal to enhance our products and solutions for the benefit of our global customer base.
Learn more: https://t.co/oRoRztWE9U
@karpathy What are your thoughts on the agent performing some form of Bayesian learning in these environments (e.g. Bayesian optimization) instead of only RL
In era of pretraining, what mattered was internet text. You'd primarily want a large, diverse, high quality collection of internet documents to learn from.
In era of supervised finetuning, it was conversations. Contract workers are hired to create answers for questions, a bit like what you'd see on Stack Overflow / Quora, or etc., but geared towards LLM use cases.
Neither of the two above are going away (imo), but in this era of reinforcement learning, it is now environments. Unlike the above, they give the LLM an opportunity to actually interact - take actions, see outcomes, etc. This means you can hope to do a lot better than statistical expert imitation. And they can be used both for model training and evaluation. But just like before, the core problem now is needing a large, diverse, high quality set of environments, as exercises for the LLM to practice against.
In some ways, I'm reminded of OpenAI's very first project (gym), which was exactly a framework hoping to build a large collection of environments in the same schema, but this was way before LLMs. So the environments were simple academic control tasks of the time, like cartpole, ATARI, etc. The @PrimeIntellect environments hub (and the `verifiers` repo on GitHub) builds the modernized version specifically targeting LLMs, and it's a great effort/idea. I pitched that someone build something like it earlier this year:
https://t.co/ANHhasxzD8
Environments have the property that once the skeleton of the framework is in place, in principle the community / industry can parallelize across many different domains, which is exciting.
Final thought - personally and long-term, I am bullish on environments and agentic interactions but I am bearish on reinforcement learning specifically. I think that reward functions are super sus, and I think humans don't use RL to learn (maybe they do for some motor tasks etc, but not intellectual problem solving tasks). Humans use different learning paradigms that are significantly more powerful and sample efficient and that haven't been properly invented and scaled yet, though early sketches and ideas exist (as just one example, the idea of "system prompt learning", moving the update to tokens/contexts not weights and optionally distilling to weights as a separate process a bit like sleep does).
"Every time she hears a rocket she starts trembling and crying. I used to tell her: 'Don’t worry. They’re not targeting us.' It’s a myth that all of us in Gaza tell our children. But it doesn't work anymore; she knows that it's a lie."
Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise
“In our simulated testing settings, with some model-external safeguards disabled, we found OpenAI's o3 and o4-mini reasoning models to be aligned as well or better than our own models”
DEEP THINK WITH CONFIDENCE
DeepConf leverages model-internal confidence signals to dynamically filter out low-quality reasoning traces
during or after generation…on AIME 2025, DeepConf@512 achieves up to 99.9% accuracy and reduces generated tokens by up to 84.7%