Can sparse training achieve wall-clock time speed up on GPU?
Yes! Simple and static #sparsity -> 2.5x faster🚀 training MLP-Mixer, ViT, and GPT-2 medium from scratch with NO drop in accuracy.
https://t.co/zpbLkiWk0o (#NeurIPS2021)
https://t.co/QbPtFEiGSy [1/6]
I finally go back to the Simons Institute for the reunion of the Lower Bounds Program! Cannot wait to learn more about GCT!
Also, stay tuned for my talk tomorrow!
Hoping to read new papers by Allen-Zhu et al. Training provably converges on greatly overparametrized deep nets. And such overparametrized deep nets can generalize when trained on data from teacher net. https://t.co/hv5DAxsApp and https://t.co/kmns5fQ43w