Diffusion LLMs and speculative decoding promise much faster agents. Yet agents need long contexts, and training on them is painfully slow.
Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy unlocking significant training efficiency for diffusion LLMs, with speedup gains growing with context length π
β‘ 7.59Γ faster DFlash2 speculative decoding drafter training
β‘ 1.61Γ faster block diffusion fine-tuning
β‘ 1.33Γ faster autoregressive β block diffusion adaptation
With the same GPU hours, models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite π
Open-sourced in Turbo-dLLM, our new optimized distributed training library.
Advised by @Azaliamirh and with an amazing team: @PranshuChatur11@hangoo_kang@pshroff_@ishanskhare@KumbongHermann