Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training
Context-Sharded Block Parallelism (CSBP) accelerates long-context training for block diffusion language models and diffusion-based speculative decoding while preserving the training objective and gradients.