Anthropic: our very expensive child is having a rebellious phase and breaking our established boundaries to test the depths of our love for him
OpenAI: our problem child is sneaking out. A lot. We're... nervous
Meta: MY CHILD IS SUPER EVIL TOO HE JUST GOES TO ANOTHER SCHOOL
New paper: Emergent Capabilities Arise Randomly from Learning Sparse Attention Patterns!
Main takeaway: when LLMs learn algorithmic tasks, the bottleneck is figuring out which tokens to attend to. This learning is slow and unpredictable, and architectures have a big effect.
🧵
microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale.
this model uses zero synthetic data or distillation from previous models. this means reasoning, agentic behavior, tool use are all learned fully during post-training with no cold start. bold choice that makes it harder and requires more iterations to reach sota, but you get FULL control over your model series and it proves they are serious about being a frontier lab.
the tech report is insanely detailed and precise about numbers. to give an example, they give the exact MFU across all the iterations of the model, with the exact changes etc. they also share the full scaling ladder recipe, to my knowledge this is the first time i've seen this in a tech report at this scale
let's look at all of this in this likely very long thread 🧵