ML ENERGY x NVIDIA blog post is live!
Power is a binding constraint for AI factories, and energy efficiency converts directly to revenue.
This collaborative blog post covers our works on providing energy observability and improving performance per watt:
- Understanding and optimizing inference energy consumption with the ML ENERGY Benchmark & Leaderboard
- Optimizing training energy consumption with Perseus & Kareus
Our works are open-source on GitHub, and we look forward to more collaborations with NVIDIA!
https://t.co/5WBSQGsr0b
@NVIDIAAI@NVIDIAAIInfra
Energy and power are first-class resources in scaling AI compute.
https://t.co/uq5rWvY6jB builds open-source infrastructure for measuring, understanding, and optimizing the energy use of ML workloads.
Start here:
- https://t.co/dlMO31ZCLT
- https://t.co/51t3ymSyEX
🚀🚀🚀Excited to share the amazing news that our work Lotto has been accepted to @USENIXSecurity 2024! Huge thanks and congratulations to my incredibly talented colleagues Peng Ye and @_ShiqiHe_, and inspiring research mentors Wei Wang, @ruichuan, and Bo Li.👏👏👏 @HkustSc (1/3)
Introducing DeepSpeed-FastGen 🚀
Serve LLMs and generative AI models with
- 2.3x higher throughput
- 2x lower average latency
- 4x lower tail latency
w. Dynamic SplitFuse batching
Auto TP, load balancing w. perfect linear scaling, plus easy-to-use API
https://t.co/iizM71bjqj