After 6+ months in the making and burning over a year of GPU compute time, we're super excited to finally release the "Ultra-Scale Playbook"
Check it out here: https://t.co/dekxY4BQZO
A free, open-source, book to learn everything about 5D parallelism, ZeRO, fast CUDA kernels, how and why overlap compute & communication – all scaling bottlenecks and tools introduced with motivation, theory, interactive plots from our 4000+ scaling experiments and even NotebookLM podcasters to tag along with you.
- How was DeepSeek trained for $5M only?
- Why did Mistral trained an MoE?
- Why is PyTorch native Data Parallelism implementation so complex under the hood?
- What are all the parallelism techniques and why were they invented?
- Should I use ZeRO-3 or Pipeline Parallelism when scaling and what's the story behind both techniques?
- What is this Context Parallelism that Meta used to train Llama 3? Is it different from Sequence Parallelism?
- What is FP8? how does it compares to BF16?
In this book, our goal was to gather, in a single place, a coherent, easy to read yet detailed story of all the techniques that make today's LLM scaling possible.
The largest factor for democratizing AI will always be teaching everyone how to build AI and in particular how to create, train and fine-tune high performance models. In other word making accessible to everybody the techniques that power all recent large language models and efficient training is possibly one of the most essential of them.
What started as a simple blog-post ended up becoming an interactive writing piece containing 30k+ words. So we've decided to actually print it as a real 100-pages physical book as well: the physical ultrafast playbook –containing all the science of distributed and fast AI training.
We plan to send free copies as gifts to the first readers of the online version so feel free to add your email in the form linked in the blog post.
This is @sayashk calmly dismantling AI scaling laws hype, during our discussion of his article he published with @random_walker earlier today. This interview slapped. #ICML2024 is a wrap!
This is precisely what statisticians used to say. "If you can predict well the next data point, they you understand the reality or the process that generates that point". They said it and said it until Causal Inference proved them wrong. Understanding takes more than prediction. #Bookofwhy
Proud to share our 150-page "proto-book" with @mmbronstein@joanbruna@TacoCohen on geometric DL! Through the lens of symmetries and invariances, we attempt to distill "all you need to build the architectures that are all you need".
https://t.co/CBN0IG8BXR
More info below! 🧵
Gershgorin circle theorem locates the eigenvalues of a matrix in the union of disks centered at the diagonal entries. Implies invisibility of diagonally dominant matrices. https://t.co/XSvyXCzuqI
It is not about chocolate, but I hope you will find it comforting. I am writing a book about learning theory. See more details and a draft in this month blog post.
https://t.co/EyCA63LfFH
An introductory and short survey on nonconvex optimization for machine learning problems https://t.co/44huz2z2PZ. A chapter of Beyond the Worst-Case Analysis of Algorithms edited by @algo_class.
Why is contrastive representation learning so useful? Our latest work shows that contrastive learning can invert the data generating process, and it paves the way towards more effective contrastive losses.
Paper: https://t.co/uhVkSTDuP7
Website/Code: https://t.co/lzaLdT7n6t
[1/5]
I’m excited to announce a remote Deep Learning Theory Summer School “at Princeton” July 27 - Aug 4 2021 aimed at grad students in math, physics, stats, cs, ee. Courses by Misha Belkin, @Andrea__M, @danintheory, Sho Yaida. More at https://t.co/hhWBcY6TMj. @david_rolnick@pfau
Q. Can we solve learning dynamics of modern deep learning models trained on large datasets?
A. Yes, by combining symmetry and modified equation analysis!
co-led with @KuninDaniel (now on twitter)
& @jvrsgsty@SuryaGanguli@dyamins
Neural Mechanics https://t.co/S8BqePyIxC
1/8
It's alarming that NeurIPS papers are being rejected based on "ethics reviews". How do we guard against ideological biases in such reviews? Since when are scientific conferences in the business of policing the perceived ethics of technical papers?
🔥Novel Video🔥The brain🧠might be doing something like Backprop after all! This paper shows that biologically plausible Predictive Coding approximates Backprop w/ only local update rules!💪
https://t.co/DxJCtQpdYY
@BerenMillidge@a_tschantz@drclbuckley@EdinburghUni@SussexUni