The project has everything you need to build your own ChatGPT clone.
🟢 What it has?
> • tokenizer (written in Rust)
> • pretraining
> • SFT (supervised fine-tuning)
> • RL (reinforcement learning)
> • model evaluation (eval)
A total of 8,000 lines of code, with no unnecessary dependencies - a perfect educational example to understand how large language model training really works.
💡 This is a project from his upcoming new course LLM101n, and a great opportunity to boost your ML skills in practice.
You can rent a GPU in the cloud and run everything yourself - the code is already ready to launch.
If you run training of the nanochat model on a cloud GPU server (for example, 8×H100), then after about 12 hours of training (cost ~ $300–400) the model reaches GPT-2 level quality on test sets (CORE-score).
And if you train for about 40 hours (cost ~ $1000), it solves simple math and coding tasks, achieving:
- 40+ on MMLU
- 70+ on ARC-Easy
- 20+ on GSM8K
🧠 This is free top-level practice from a master that you shouldn’t miss.
GitHub
#LLM #nanochat #MachineLearning #DeepLearning #AI #GPT
🤖 Data Science, ML & Big Data with @DataXplore
