🟣 How it was made?
☞ First it was transformed from a pre-trained autoregressive model and then trained on 500 billion tokens to completely change behavior of diffusion model.
☞ Regular models (AR, autoregressive) write text word by word, while RND1 generates the entire sentence at once and then refines it step by step, as if "developing" the text from noise.
☞ This is a Diffusion Language Model (DLM), an analogue of diffusion models that create images, but here it "draws" words.
☞ The Radical Numerics team figured out how to turn a ready model into a diffusion one without training from scratch.
☞ They simply changed the attention type and fine-tuned the model on a new task.
☞ This method is called AR-to-Diffusion Conversion (A2D) that is, conversion from an autoregressive model to a diffusion model.
🟠 How it works?
1️⃣ Take a strong GPT-like model.
2️⃣ Change attention mechanism, now the model sees the entire context at once.
3️⃣ Continue training on the diffusion task.
4️⃣ Use different learning rates for different parts of the network so that the model doesn’t forget the old but learns a new way of thinking.
Under the hood
☞ Mixture-of-Experts (MoE) : the model has 30 billion parameters, but only 3 billion are active at a time. This makes it powerful yet efficient.
☞ Continuous fine-tuning : old knowledge is not erased but "embedded" into the new mode.
☞ Huge batches : the model learns on large data batches to stabilize training since it doesn’t process all tokens at once.
🟢 What makes RND1 more interesting?
☞ Parallel generation : text is created faster without step-by-step delay.
☞ Lower costs : only 3 billion active parameters, yet quality comparable to large GPTs.
☞ New architecture : opens the way for hybrid models combining the advantages of AR and DLM.
☞ Fully open code and weights : you can explore, modify, and run it yourself.
☞ The first serious step toward self-improving AI : the model can not only learn but also help design the next version.
This is a truly interesting method; RND1 shows that AI can be not just trained but restructured — changing its very logic of thinking without starting "from scratch."
It seems this could become the foundation for Recursive Self-Improvement (RSI) systems, where AI is capable of creating and improving itself.
Code, Report, Weights
#RND1 #RadicalNumerics #DLM #DiffusionModel #MoE #OpenSource
🤖 Data Science, ML & Big Data with @DataXplore