TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 585 subscribers
Post #1726 133
RND1 - new experimental model with 30 billion parameters built on Sparse Mixture-of-Experts system with 3 billion active parameters.

🟣 How it was made?
☞ First it was transformed from a pre-trained autoregressive model and then trained on 500 billion tokens to completely change behavior of diffusion model.

☞ Regular models (AR, autoregressive) write text word by word, while RND1 generates the entire sentence at once and then refines it step by step, as if "developing" the text from noise.

☞ This is a Diffusion Language Model (DLM), an analogue of diffusion models that create images, but here it "draws" words.

☞ The Radical Numerics team figured out how to turn a ready model into a diffusion one without training from scratch.

☞ They simply changed the attention type and fine-tuned the model on a new task.

☞ This method is called AR-to-Diffusion Conversion (A2D) that is, conversion from an autoregressive model to a diffusion model.


🟠 How it works?
1️⃣ Take a strong GPT-like model.
2️⃣ Change attention mechanism, now the model sees the entire context at once.
3️⃣ Continue training on the diffusion task.
4️⃣ Use different learning rates for different parts of the network so that the model doesn’t forget the old but learns a new way of thinking.

Under the hood

☞ Mixture-of-Experts (MoE) : the model has 30 billion parameters, but only 3 billion are active at a time. This makes it powerful yet efficient.

☞ Continuous fine-tuning : old knowledge is not erased but "embedded" into the new mode.

☞ Huge batches : the model learns on large data batches to stabilize training since it doesn’t process all tokens at once.


🟢 What makes RND1 more interesting?
☞ Parallel generation : text is created faster without step-by-step delay.
☞ Lower costs : only 3 billion active parameters, yet quality comparable to large GPTs.
☞ New architecture : opens the way for hybrid models combining the advantages of AR and DLM.
☞ Fully open code and weights : you can explore, modify, and run it yourself.
☞ The first serious step toward self-improving AI : the model can not only learn but also help design the next version.

This is a truly interesting method; RND1 shows that AI can be not just trained but restructured — changing its very logic of thinking without starting "from scratch."

It seems this could become the foundation for Recursive Self-Improvement (RSI) systems, where AI is capable of creating and improving itself.


Code, Report, Weights

#RND1 #RadicalNumerics #DLM #DiffusionModel #MoE #OpenSource

🤖 Data Science, ML & Big Data with @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →