This language model (~1B parameters) optimized for efficient *on-device* operation.
🟢 Why it's good?
Outperforms Gemma 3 1B and Llama 3.2 1B in reasoning, knowledge, and long context tasks, supporting up to 128,000 tokens.
Thanks to hybrid attention (local + global in a 3:1 ratio, window 512), it achieves low latency and KV-cache memory savings.
Quantization to 4-bit (int4) barely reduces quality:
• CPU - group weight quantization and dynamic activation
• GPU - per-channel quantization
The model has also undergone instruction fine-tuning, making it suitable for communication, generation, and text processing tasks.
Model
🤖 Data Science, ML & Big Data with @DataXplore