Excited to share Context-Sharded Block Parallelism (CSBP) and Turbo-dLLM, our new optimized distributed training library for Diffusion LLMs!
Our distributed parallelism strategy unlocks significant training efficiency for diffusion LLMs, with speedup gains growing with context length ๐
On 8x H100 GPUs, Turbo-dLLM accelerates DFlash2 speculative-decoder training by 2.48x at 512K and 7.59x at 1M context length.
Simply pip install turbo-dllmโก๏ธ or check out https://t.co/Mcme2NOA3Y.
Honored to work with @TarunSures41845, @PranshuChatur11, @hangoo_kang, @pshroff_ , @KumbongHermann, and advisor @Azaliamirh ๐!