So you want to skip our thinning proofs—but you’d still like our out-of-the-box attention speedups? I’ll be presenting the Thinformer in two ICML workshop posters tomorrow!
Catch me at Es-FoMo (1-2:30, East hall A) and at LCFM (10:45-11:30 & 3:30-4:30, West 202-204)
Your data is low-rank, so stop wasting compute! In our new paper on low-rank thinning, we share one weird trick to speed up Transformer inference, SGD training, and hypothesis testing at scale. Come by ICML poster W-1012 Tuesday at 4:30!
If you’re not at ICML, don’t worry! You can still read our work. Our new theoretically principled algorithms beat recent baselines across multiple tasks—including Transformer approximation! https://t.co/GZphBh4L83
Your data is low-rank, so stop wasting compute! In our new paper on low-rank thinning, we share one weird trick to speed up Transformer inference, SGD training, and hypothesis testing at scale. Come by ICML poster W-1012 Tuesday at 4:30!
At ICML this week?
Check out @annabelle_cs's paper in collaboration with @LesterMackey and colleagues on Low-Rank Thinning!
⏰ Tue 15 Jul 4:30 - 7 p.m. PDT
New theory, dataset compression, efficient attention and more: https://t.co/Wc1DUrHp7z
GRAD SCHOOL APPLICATION(2.0) 🧵
Got multiple fully funded PhD offers recently and realized from conversations I have been having that many people don't approach the application process intentionally.
Sharing my application process doc as an example below.
Open and Retweet 🔃
@boazbaraktcs@tomgoldsteincs I am wondering if they test copyright protection-enhanced models alongside the vanilla generative models, to see if your provable method works empirically at scale! IIRC you tested empirically on CIFAR to verify your method qualitatively.
@tomgoldsteincs It might have come out during/after your work, but have you all compared to Nikhil Vyas @boazbaraktcs's work on copyright protection? https://t.co/hV7zVVjm2m
Optimizer tuning can be manual and resource-intensive. Can we learn the best optimizer automatically with guarantees?
With @HazanPrinceton, we give new provable methods for learning optimizers using a control approach. Excited about this result!
https://t.co/GTpNSdcQlm
(1/n)
Neural networks are non-convex, and non-smooth. Unfortunately, most theoretical analysis is either convex, or smooth. Should we abandon the past? No! With @bremen79 and @n0royalroad, we import prior know-how via an *online to non-convex* conversion: https://t.co/EzbaXO8dtT.
My favorite non-ML paper I read this year is probably "Bayesian Persuasion" (2011), which I somehow only found out about recently. Simple & beautiful.
The first 2 pages are sufficient to be persuaded.
https://t.co/HwYfiCmN5u
In the LLM-science discussion, I see a common misconception that science is a thing you do and that writing about it is separate and can be automated. I’ve written over 300 scientific papers and can assure you that science writing can’t be separated from science doing. Why? 1/18
1/Is scale all you need for AGI?(unlikely).But our new paper "Beyond neural scaling laws:beating power law scaling via data pruning" shows how to achieve much superior exponential decay of error with dataset size rather than slow power law neural scaling https://t.co/Vn62UJXGTd
New ICML workshop paper 🚨https://t.co/5wR1ZqbOyj. Are deep neural nets calibrated? The literature is conflicted... bc the question itself has changed over time: as architectures, optimizers, and datasets evolve, it is difficult to disentangle factors which affect calibration.