Introducing Domino: a novel zero-cost communication tensor parallelism (TP) training engine for both single node and multi-node settings.
- Near-complete communication hiding
- Novel multi-node scalable TP solution
Blog: https://t.co/08bPanyr9M
1/ Announcing the development of OpenMoE project! 🚀
Open Mixture-of-Experts Language Models!
MoE + UL2 objective + umT5 tokenizer + 50% code data mix.
GitHub: https://t.co/dxcMsFRKY3
Blog: https://t.co/AznT0PfdFE
I’m attending #ISCA co-located with #FCRC 🎉 We will present two papers at the MlArchSys and ASSYST workshop on #LLM and #GAN at Canary 2. Feel free to drop by and say hi! @ISCAConfOrg
(2/2) We warmly welcome all professionals in the field to join us, engage in enriching conversations, and contribute to our vibrant community! #LLM#Singapore#MLSys#AI#CommunityBuilding 🤝
(1/2) As large-scale models continue to evolve, the need for associated foundational systems is also growing. We've set up an MLSys discussion group (https://t.co/R5qcTlglXM), planning to host bi-weekly discussions on academic papers or updates on cutting-edge advancements.
(1/2) As large-scale models continue to evolve, the need for associated foundational systems is also growing. We've set up an MLSys discussion group (https://t.co/R5qcTlglXM), planning to host bi-weekly discussions on academic papers or updates on cutting-edge advancements.
🧵🧠 We're witnessing incredible scientific progress in image & text reconstruction from fMRI nowadays. But what about reconstructing video from fMRI? Allow me to introduce our recent preprint: Mind-Video
https://t.co/VL2KXz8o9K
https://t.co/KyNtsCxDIJ
https://t.co/bhjz0PDlS6
@XueFz@Francis_YAO_ From the HW point of view, I think sparse computation ( eg. cuSPARSE) can be used to accelerate the sparse activations, while the MoE models may not, as they operate on a much coarser level. But some architectural changes may help (eg load balancing among experts etc).
@YisongMiao Thanks! The DAG representation of NN was originally proposed by others (eg. #TensorFlow ), but we leverage it to shrink the search space for parallel strategy into a much smaller domain, where we are more comfortable optimizing. Hopefully, it paves the way for #LLM research :)
TAP: Accelerating Large-Scale DNN Training Through Tensor Automatic Parallelisation. Ziji Shi, Le Jiang, Ang Wang, Jie Zhang, Xianyan Jia, Yong Li, Chencan Wu, Jialin Li, and Wei Lin https://t.co/d4BnMvXMBT
We recently uploaded our work with Alibaba on *quickly* and *automatically* finding the optimal #tensorparallel strategy for #LLM. Compared to SoTA approaches, we are ~20-160x faster. Comments are welcomed!
Arxiv: https://t.co/mxEhMrKF8R
#ChatGPT has been phenomenal, but have you ever wondered how it was trained? In fact, finding the optimal parallel strategy for such LLM is very challenging, as the candidate space grows exponentially w.r.t size. (1/2)
#ChatGPT has been phenomenal, but have you ever wondered how it was trained? In fact, finding the optimal parallel strategy for such LLM is very challenging, as the candidate space grows exponentially w.r.t size. (1/2)