🗳️ LLM101n: 从零开始实现大型语言模型
🌐 Karpathy 开发的LLM教程
LLM101n是Andrej Karpathy开发的一套从零开始实现大型语言模型(LLM)的教程和代码库。在本课程中,我们将构建一个 Storyteller AI 大语言模型 (LLM)。携手并进,您将能够使用 AI 创建、完善和阐释小故事。我们将使用 Python、C 和 CUDA 从头开始,以最少的计算机科学先决条件构建从基础知识到类似于 ChatGPT 的功能性 Web 应用程序的端到端的一切。最后,您应该对 AI、LLMs 和更广泛的深度学习有相对深入的了解。
Syllabus:
Chapter 01 Bigram Language Model (language modeling)
Chapter 02 Micrograd (machine learning, backpropagation)
Chapter 03 N-gram model (multi-layer perceptron, matmul, gelu)
Chapter 04 Attention (attention, softmax, positional encoder)
Chapter 05 Transformer (transformer, residual, layernorm, GPT-2)
Chapter 06 Tokenization (minBPE, byte pair encoding)
Chapter 07 Optimization (initialization, optimization, AdamW)
Chapter 08 Need for Speed I: Device (device, CPU, GPU, ...)
Chapter 09 Need for Speed II: Precision (mixed precision training, fp16, bf16, fp8, ...)
Chapter 10 Need for Speed III: Distributed (distributed optimization, DDP, ZeRO)
Chapter 11 Datasets (datasets, data loading, synthetic data generation)
Chapter 12 Inference I: kv-cache (kv-cache)
Chapter 13 Inference II: Quantization (quantization)
Chapter 14 Finetuning I: SFT (supervised finetuning SFT, PEFT, LoRA, chat)
Chapter 15 Finetuning II: RL (reinforcement learning, RLHF, PPO, DPO)
Chapter 16 Deployment (API, web app)
Chapter 17 Multimodal (VQVAE, diffusion transformer)
Further topics to work into the progression above:
Programming languages: Assembly, C, Python
Data types: Integer, Float, String (ASCII, Unicode, UTF-8)
Tensor: shapes, views, strides, contiguous, ...
Deep Learning frameowrks: PyTorch, JAX
Neural Net Architecture: GPT (1,2,3,4), Llama (RoPE, RMSNorm, GQA), MoE, ...
Multimodal: Images, Audio, Video, VQVAE, VQGAN, diffusion
🔥Karpathy昨天刚上传的github repo,内容尚未更新,可以先收藏下来!!!
Related Links:
- LLM101n GitHub仓库
🏷️ #course
🐟 @protoolkit
All about AI and Productivity.
Post #90
564
- 👍 1