☞ RoPE with YaRN + NTK-by-parts for context scaling
☞ RMSNorm
☞ SwiGLU with clamping and residual connections
☞ Mixture-of-Experts (MoE)
☞ Self-Attention, optimized via Grouped Query Attention (GQA)
☞ Learned sinks
☞ Banded (sliding window) attention
☞ Support for KV caching
All of this works on a single A100 SXM (80GB). He also wrote detailed documentation with the theory of each component, as well as instructions for setup and inference.
Repository
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
