Remember the slime project? a lightweight tool used in many modern post-training pipelines.
Slime proved that a lightweight design works, and Miles takes the next step like large-scale training of MoE architectures and support for heavy industrial workloads.
🟢 What Features+Stability Miles offers?
Miles offers what is called "True On-Policy." Previously, there was often a divergence between training and inference. Now, thanks to an infrastructural approach, LMSYS has achieved zero divergence. This was made possible by using Flash Attention 3, the DeepGEMM library, and kernels from Thinking Machines Lab, working in conjunction with torch.compile.
The second feature is the use of speculative decoding. Usually, in RL, the draft model is frozen, which prevents it from following the target model's policy. LMSYS added online training of the draft model.
Test results are positive: generation speedup of more than 25%, especially in the later stages of training.
STABILITY.
For enterprise, memory is money. Miles include mechanisms that prevent system crashes due to non-critical OOM errors and fix excessive memory consumption in FSDP.
The project promises support for multimodal training, compatibility with SGLang v2 and extended speculative decoding.
Github #AI #ML #RL #Miles #LMSYS
••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
