TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #1811 230
How to effectively train a model even on consumer GPUs?

Two sequence parallelism methods are combined In ModelScope SWIFT.

🟢 Ulysses + Ring Attention
✅ Ulysses - splits attention by heads, uses almost no traffic (but limited by the number of heads)

✅ Ring Attention - scales beyond the number of heads through ring P2P communications, with "zig-zag" balancing for causal models

Ulysses works first, and only when it can no longer handle the load (e.g., GQA or cluster >8 GPUs), Ring is activated.

Sequence split is built directly into model's forward-hook : no data hacks, full compatibility with FlashAttention.

Result on Qwen2.5-3B with 65k tokens:
75.4 GiB → 17.9 GiB VRAM on 8× A100
Works with SFT, DPO, GRPO, multimodality, and padding-free inputs.


Enabled with a single flag command:
--sequence_parallel_size 8

More details on GitHub

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →