In the fight to defend openness in AI, the Marin project is a precious demonstration of openness in model training, with open code, data, recipes, even experimental results. Releasing AI research openly used to be the norm; I'm grateful for @percyliang's open lab approach.
At Marin we’ve found that scaling laws let us simulate the entire training trajectory. This became clear in our last 67B MoE run. The run hit the loss target within 1%, but perhaps more interestingly, arbitrary points during training could also be predicted to within 1%. This model is finishing up long context and will soon enter RL. Based on this finding, we’ve used a scaling ladder to simulate the training trajectory of the 535B MoE we kicked off yesterday. Will see how it goes. Now if someone can figure out how to fit a scaling law to each parameter we can just predict our final model and skip this training business.
Something else I've learned from this process is that getting a training run off the ground is way more than just fiddling with the architecture. Lot of herculean efforts from teammates to get trillions of tokens curated, get hardware running on a new GPU stack, improve MFU, and propose great ideas that were tested and integrated.
Given that this is our largest run to date, I anticipate lots of learning and adapting along the way. But thats part of the fun. Details: https://t.co/N7ZxcgvkKs
🧬🧬🧬🧬🧬🧬 huge drop for DNA models!!
Variant effect prediction is probably the most important task for DNA foundation models because it directly effects human health outcomes.
And so critical that this team is not trying to reinvent the wheel on the backbone for this model, which means that it will run fast, compatibly, and be easy to develop futher.
You see a lot of people in AI for science making the mistake of trying to do architecture improvements or fundamental changes, which usually means the model doesn't work nearly as well and takes way longer to develop. Such an exemplar release tbh
Excited to share MarinDNA, a 1B gLM that rivals Evo 2 40B on variant effect prediction while being 2,330x faster.
With @eczech0, we built around a standard Transformer so we could reuse LLM infra and methods while focusing on data curation and scaling.
https://t.co/TVChdjJyMP 🧵
🚨Coming to #ICML2026 🇰🇷
Last year, we introduced Block Diffusion LMs, adopted by dLLMs at Google, NVIDIA, Meta, etc.
🚀Today, we introduce Set Diffusion, a faster, more flexible alternative that combines AR and diffusion by interpolating generation orders.
🧵1/8
New OSS diffusion LLM out of a major lab, this time @nvidia.
Shout-out to Cornell students @mariannearr and @SchiffYair whose work on block diffusion and encoder-decoder diffusion is at the core of this model along with MDLM.
Marianne's and Yair's papers are linked below.
Building momentum at Marin! Upgrading from Dense -> 129B parameter MoEs -> architecture improvements -> optimizer improvements gives our pretraining recipe an estimated 6x cumulative learning speedup, accounting for MFU. Includes community contributions. https://t.co/5dPB9uBiSp
Not only do we want to train a good model, we want to know it'll be good before we even start training.
About a month ago, the Marin team launched a 129B (16B active) 1e23 FLOPs MoE run and preregistered a loss of 2.252. The run finished this past week and landed at 2.234.
https://t.co/OptaVa7jIO
To train better open models, we need predictable scaling.
Delphi is Marin’s first step: we pretrained many small models with one recipe, then extrapolated 300× to predict a 25B-param / 600B-token run with just 0.2% error.
Getting there took some work 🧵
I've written up a new post, trying to distill the spirit of this essay into something much shorter. It is about the assumptions that interpretability researchers make, and what we might hope to learn about ourselves from studying neural networks:
https://t.co/0yhd9FXuBs
Researchers' brilliant ideas often get lost in the sea of endless SOTA claims on weak baselines. At Marin we battle-test ideas in an open arena, where anyone's idea can be promoted to the next hero run. One that recently rose up was @Jianlin_S MoE Quantile Balancing, used in our last 1e22 and ongoing 130B run. Animated visuals of how QB performed are available in the OpenAthena blog. https://t.co/BDSsonuNH7
after starting open Athena less than two years ago, it's amazing to see the progress to advance ai for science that the team has driven https://t.co/qeO2nHT8Iz
How far do Marin's scaling laws extrapolate? At least 100x, apparently!
Despite spooky spikes, our 1e23 Delphi finished on forecast. The compute-optimal ladder costs ~1e21 FLOPs to train. Good scaling science lets you “run” this (not tiny) experiment at 1/100th the cost.
Recently had the pleasure of lecturing back at Princeton in a grad seminar. I took the opportunity to cover how scaling laws have evolved since their inception, leaning heavily on great external content from my colleagues @borgeaud_s@jalayrac@jacobaustin132 .
Content in thread
Our 1e23 "Delphi" (~25B param model trained for ~600B tokens) run for Marin has entered its learning rate decay phase.
Lots of spikes at this scale, very scary! Despite that, the run is looking on track to be close to our pre-registered scaling laws predictions. Stay tuned...
Optimization theory for adaptive methods actually predicts most of what we know about hyperparameter scaling in LLM pretraining, and suggests new strategies as well. We did a deep dive here.
I have kids. I work in AI every day. And honestly? I have no idea what their careers will look like in 15 years. But I know what will carry them through.
First, and this might sound unromantic: make money and save it for them. We can debate educational philosophy all day, but the world is changing so fast that financial security might be the most practical gift we can give. Buy some gold bars. Seriously.
Second, nurture their imagination. AI rewards people with initiative and wild ideas. The kid who daydreams, who asks weird questions, who wants to try ten things at once? That kid will thrive. AI can execute. AI can be disciplined. What AI can't do is dream up something nobody's thought of before.
Third, build resilience. There are no more iron rice bowls (guaranteed lifetime jobs). Any stable, predictable job is exactly the kind of job AI will learn to replace. Our kids will likely switch directions many times in their lives. Learn something new, get replaced, pivot, repeat. It's more like being a hunter than a farmer. Schools don't teach this. Schools teach you to follow a linear path: high school, college, grad school, stable job. That linear path is becoming the most dangerous one.
Last, invest in their ability to connect with other humans. Not networking. Not schmoozing. Real emotional connection. Building trust, offering support, making people feel seen. As AI handles more of the rational, analytical work, the human ability to genuinely relate to other humans becomes more rare and more valuable.
I don't have all the answers. But I know that imagination, resilience, and genuine human warmth aren't going out of style anytime soon.
#AI #Parenting #Education #FutureOfWork
📉📉NEW SCALING LAW PHENOMENON 📉📉
We find that knowledge and reasoning exhibit different scaling behaviors!
Super excited to finally tell you all about our paper on the compute optimal scaling of skills:
https://t.co/SH3YCMyIeG
[1/n]