斯坦福提出新并行策略提升扩散模型训练效率
Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy enablin...
斯坦福团队新提出的CSBP并行策略,能让扩散模型训练速度大幅提升,特别是长上下文场景,值得看看。
斯坦福提出的Context-Sharded Block Parallelism (CSBP)策略让扩散模型训练速度提升显著,在DFlash2训练中达到7.59倍加速,在块扩散微调中达到1.61倍加速。使用该策略训练的模型在SWE-bench和Terminal-Bench基准测试中表现更好。
Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy enablin...
Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy enabling significant training efficiency for diffusion LLMs, with gains that grow with context length 🚀 ⚡ 7.59× faster DFlash2 speculative decoding drafter training ⚡ 1.61× faster block diffusion fine-tuning ⚡ 1.33× faster autoregressive → block diffusion adaptation Tarun Suresh @ COLM 2026 @TarunSures41845 Diffusion LLMs and speculative decoding promise much faster agents. Yet agents need long contexts, and training on them is painfully slow. Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy unlocking significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀 ⚡ 7.59× faster DFlash2 speculative decoding drafter training ⚡ 1.61× faster block diffusion fine-tuning ⚡ 1.33× faster autoregressive → block diffusion adaptation With the same GPU hours, models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite 👑 Open-sourced in Turbo-dLLM, our new optimized distributed training library. Advised by @Azaliamirh and with an amazing team: @PranshuChatur11 @hangoo_kang @pshroff_ @ishanskhare @KumbongHermann 🔗 View Quoted Tweet 💬 1 🔄 2 ❤️ 19 👀 2013 📊 3 ⚡