Карпатый зарелизил новый репозиторий для обучения LLM’ок с нуля — 🔗nanochat
It weighs ~8,000 lines of imo quite clean code to:
- Train the tokenizer using a new Rust implementation
- Pretrain a Transformer LLM on FineWeb, evaluate CORE score across a number of metrics
- Midtrain on user-assistant conversations from SmolTalk, multiple choice questions, tool use.
- SFT, evaluate the chat model on world knowledge multiple choice (ARC-E/C, MMLU), math (GSM8K), code (HumanEval)
- RL the model optionally on GSM8K with "GRPO"
- Efficient inference the model in an Engine with KV cache, simple prefill/decode, tool use (Python interpreter in a lightweight sandbox), talk to it over CLI or ChatGPT-like WebUI.
- Write a single markdown report card, summarizing and gamifying the whole thing.
С какими еще техниками можно тут поупражняться:
⏺Архитектура — упрощенная Llama
⏺Rotary Embeddings, QK Normalization, Multi-Query Attention (MQA)
⏺ReLU² активации, logit softcapping
⏺Muon+AdamW из modded-nanoGPT
Ресурсы: ~4 часов 8XH100
В общем, берем на заметку🙌
🔗https://github.com/karpathy/nanochat