Excited to share Context-Sharded Block Parallelism (CSBP) and Turbo-dLLM, our new optimized distributed training library for Diffusion LLMs!
Our distributed parallelism strategy unlocks significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀
On 8x H100 GPUs, Turbo-dLLM accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M context length.
Simply pip install turbo-dllm⚡️ or check out https://t.co/Mcme2NOA3Y.
Honored to work with @TarunSures41845, @PranshuChatur11, @hangoo_kang, @pshroff_ , @KumbongHermann, and advisor @Azaliamirh 🌟!
@thewilliampan i fully agree but counter: if you advertise only what you know and nothing more, it is much harder to get your foot in the door to somewhere (eg. groundbreaking companies and startups) where you can actually gain the knowledge and experience to walk the walk
@AllenNaliath@emilyzsh@sama you could get it lightly engraved but the only issue is that the metal is so thin and the components underneath are so sensitive that it might cause damage. maybe finding a way to stain the metal underneath could be the move
@ShrayAlag@sahiladhawade@PhilzCoffee@siddrrsh on the other hand, @siddrrsh can flip a switch and context switch so fast that momentarily, he becomes the only person to ever achieve true multiprocessing on human hardware. hard to tell which state hes in sometimes because either way, hes wired in
@siddrrsh the issue here is even if you implement these papers, only the companies with major backing can actually train these models and get the fruits of their labor, enabling them to repeat the process; it’s a positive feedback loop that keeps the common man out of the race