Plot twist 😅: it turns out something went wrong in the HAL → arXiv transfer, and arXiv received the wrong PDF. This was outside the control of arXiv, my student, and me. Mystery solved - many thanks to arXiv for clarifying, and for all they do for the research community!
update: our arXiv rejection made it to NeurIPS!!!!
https://t.co/J9BYAgynSp
“08/26: Thank you for submitting your work to arXiv. We regret to inform you that arXiv's moderators have determined that your submission will not be accepted and made public on https://t.co/o6xqSIeSmc.”
The reason being "Our moderators determined that your submission does not contain sufficient original or substantive scholarly research and is not of interest to arXiv."
And then we attacked them with AI. More than 30 now have a complete proof or an explicit counterexample. Some have already been checked by mathematicians, a few formalized in Rocq. We welcome reviewers!
https://t.co/OLniRAqxiX
Maybe I should publish my "Hacking https://t.co/YOsRl0IpEo and not getting any reply after they fixed it right away" or even better "Hacking https://t.co/YOsRl0IpEo who barely Hacked OpenAI, after I actually hacked OpenAI." Too many hack too many open and too many ai. Probably.
Stop with the hype guys, just get to work.
We’ve been researching fault tolerance in Crucible, Templar’s pre-training platform.
The goal: keep training through node failures and make better use of unreliable workers and spot instances, without idling an entire model replica when one stage goes down.
1/n
@aidrinksyoshake@WanderSchnee@stevenstrogatz Thank you for speaking up about this. More generally, the way some theorists react to this can be pretty off-putting, as if they thought they were doing something more important.
If you're at #ICML2026, come chat about foundation models for physics!
Posters:
• Test-time generalization for physics, led by @LouisSerrano31 (Wed 8, 5pm)
• Walrus, led by @mikemccabe210 (Wed 8, 10:30am)
Both are collaborations with @PolymathicAI and the @FlatironInst
Europe plans to lead in frontier AI. Meanwhile, an ERC review about AI from a colleague:
"The proposal contains a paragraph on risks, but unfortunately does not discuss the main risk, namely this is such a fast moving research area that it is very likely that most of these ideas are currently being explored by other groups in academia and industry (possibly with some of them having access to resources well beyond an ERC grant)"
Training LLMs over low-bandwidth networks is hard but its necessary in the current landscape. In pipeline parallelism, every stage sends its activations to the next stage, and as models scale, this communication becomes the main bottleneck.
Our new work 🍁MAPL asks: Can we cut this communication sharply without hurting model quality and token efficiency ?
🧵 1/N
Training big models gets painful once a full replica won't fit on one accelerator.
You end up with model-parallel methods or techniques like FSDP that are communication-heavy and limited in how far they parallelize.
We tried a new axis that lets you split the model the way model parallelism does, but communicate gradients instead of activations.
🧵 1/N