π Global, Local, Sparse: Attention Patterns in Long-Context Transformers
The O(nΒ²) complexity of dense (global) attention is impractical for long sequences. Here's what ML engineers need to know about the three dominant patterns: π§ βοΈ
1οΈβ£ Global (Full Dense) π
β Every token attends to every token.
β A = softmax(QKα΅ / βd) V
β Complexity: O(nΒ²d)
β Use: Short contexts (<4k) or precise recall tasks. π―
β Downside: KV cache memory explodes. π₯
2οΈβ£ Local (Sliding Window) β e.g., Mistral πͺ
β Tokens attend to a fixed neighborhood (Β±512).
β Complexity: O(n Β· w)
β Use: Streaming text, audio, DNA. π§π§¬
β Trade-off: Linear scaling but zero long-range mixing between windows. π
3οΈβ£ Sparse β e.g., BigBird, Longformer πΈ
β Pattern: Local + Global (e.g., [CLS] tokens) + Random/strided.
β Complexity: O(n Β· (w + g + r)) β O(n)
β Use: Document summarization (5kβ16k tokens). π
β Insight: Sparse graphs preserve universal approximation if graph diameter is bounded. π
Where we're going: Static sparsity is losing to dynamic routing (Mixture of Depths, 2024). π Also, linear RNN-like attention (Mamba, RWKV) challenges whether we need any static pattern. π€
https://t.me/MachineLearning9 π‘
Post #5887
2.41K

- β€ 8