Zyphyra with AMD & IBM, disproved the popular belief that serious neural network training is only possible on chips from one well-known company.
🟢 How experiment practically proved an alternative?
The project setting was truly "red": AMD Instinct GPUs, AMD Pensando network interfaces, and the ROCm software stack.
ZAYA1 turned out to be quite interesting. It has 8.3 billion total parameters, of which only 800 million are active.
Despite its compactness, it performs well in tests. In reasoning, mathematics, and programming, ZAYA1 outperformed Llama-3-8B and OLMoE. And overall, it stands alongside Qwen3-4B and Google's Gemma3-12B.
Training took place on an IBM Cloud cluster, where the model processed 14 trillion tokens. But it’s not just about the hardware; architectural innovations were used in the pipeline:
A new attention mechanism - Compressed Convolutional Attention. It uses convolutions inside the attention block, reducing computational and memory load.
Redesigned MoE router. Instead of the standard linear router, ZAYA1 uses a complex sequence of operations that make the "experts" inside the neural network specialize much better.
Residual Scaling. Trainable scalar gates were added to the residual stream at the outputs of each block, allowing the model to control the degree of forgetting.
To run inference, zaya branch of the transformers fork from the Zyphra repository is required.
Apache 2.0 License, Article, Model, Arxiv
#AI #ML #LLM #MoE #Zyphra
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
