Keeping corrupted K/V local halves attention traffic for NemotronDiffusion 14B, which uses full attention.
DiffusionGemma 26B-A4B saves more because 25 of its 30 layers use a 1,024-token sliding window: there, CSBP sends only
the clean rows each GPU's queries can see, while the CP8 baseline exchanges full shards in every layer.
| Measure | NemotronDiffusion 14B CP8 → CSBP | DiffusionGemma 26B-A4B CP8 → CSBP |
| Step time (s) | 66.50 → 56.37 1.18× | 49.66 → 34.13 1.45× |
| Forward GPU span (s) | 17.57 → 13.00 −26.0% | 10.99 → 5.83 −47.0% |
| Backward GPU span (s) | 49.46 → 45.02 −9.0% | 36.50 → 26.62 −27.1% |
| Attention traffic (GiB/GPU) | 210.00 → 105.00 −50.0% | 433.13 → 28.30 −93.5% |
| Exposed attention comm. (s/GPU) | 3.034 → 2.414 −20.4% | 10.939 → 0.723 −93.4% |