A powerful open Mixture-of-Experts model trained on GLM-4.5 Air Base with two stages: SFT and large-scale RL fine-tuning.
🟢 The Features:
This is the first model of this scale where Asynchronous RL is not an experiment but the foundation of training. As a result, the model demonstrates strong performance in mathematics, coding, and reasoning.
The model's focus is on long action chains and agent tasks, not just text generation.
- The model shows top results for its size in mathematics, coding, and reasoning.
- Training was conducted on 512×H200 for about ~2 months.
- Used own stack: PRIME-RL, Verifiers, Environments Hub, and sandbox infrastructure.
- Everything is open: code, environments, tools.
Technical Report, HF, PRIME-RL, Verifiers, Environments Hub
#AI #intellect3 #Primeintellect #GLM45
🤖 @DataXplore for Data Science, ML, DL, NN & Big Data Analysis