Most open models today fall into two categories: either massive and powerful, or small and efficient. Rarely both.
Sber’s R&D team released GigaChat-3.1 Ultra and Lightning under MIT, covering both ends in a single lineup. Both models are pretrained from scratch on internal infrastructure, without relying on external finetuning.
👉 Breakdown:
🧠 Ultra — 702B MoE
outperforms DeepSeek-V3-0324 and Qwen3-235B, supports FP8 and MTP, runs on 3 HGX
⚡ Lightning — 10B MoE
matches Qwen3-1.7B in speed, surpasses Qwen3-4B and Gemma-3-4B, with 256k context
Both models are multilingual (14 languages) with a focus on English and Russian. GigaChat here works as a unified foundation — scaling from local inference to high-performance systems without changing the stack.
Drop a like if you want to see more posts like this 👍❤️
Post #2279
6.75K

- ❤ 18