Two sequence parallelism methods are combined In ModelScope SWIFT.
🟢 Ulysses + Ring Attention
✅ Ulysses - splits attention by heads, uses almost no traffic (but limited by the number of heads)
✅ Ring Attention - scales beyond the number of heads through ring P2P communications, with "zig-zag" balancing for causal models
Ulysses works first, and only when it can no longer handle the load (e.g., GQA or cluster >8 GPUs), Ring is activated.
Sequence split is built directly into model's forward-hook : no data hacks, full compatibility with FlashAttention.
Result on Qwen2.5-3B with 65k tokens:
75.4 GiB → 17.9 GiB VRAM on 8× A100
Works with SFT, DPO, GRPO, multimodality, and padding-free inputs.
Enabled with a single flag command:
--sequence_parallel_size 8
More details on GitHub
🤖 Data Science, ML & Big Data with @DataXplore
