The Approach
Sebastian Raschka's LLMs-from-scratch takes the opposite approach from most LLM resources: instead of teaching you how to use models, it teaches you how to build them. Every component — tokenization, embedding, attention, transformer blocks, training loops — is implemented from scratch in PyTorch with detailed explanations.
The repo is the companion to Raschka's bestselling book Build a Large Language Model (From Scratch), but it stands on its own. Each chapter is a self-contained Jupyter notebook with complete, runnable code.
What You'll Build
The repo walks through constructing a GPT-class model step by step:
- Chapter 2 — Text data processing: tokenization, byte-pair encoding, and data loaders
- Chapter 3 — Attention mechanisms from scratch: self-attention, multi-head attention, causal masks
- Chapter 4 — Full GPT architecture: transformer blocks, layer normalization, residual connections
- Chapter 5 — Pre-training: training loops, loss functions, learning rate schedules
- Chapter 6 — Fine-tuning for classification: adapter layers and LoRA
- Chapter 7 — Instruction fine-tuning: creating a chat-capable model from a base model
Why It Matters
In an era where most developers interact with LLMs through APIs, understanding what happens inside the model is a genuine competitive advantage. When you've built attention from scratch, you understand why context windows have limits. When you've implemented a training loop, you understand why fine-tuning costs what it does. This knowledge makes you a better engineer, even if you never train a model from scratch in production.
This repo is the antidote to "API developer" syndrome. Raschka's approach is the CS degree of LLMs — it teaches you the fundamentals so deeply that everything built on top becomes intuitive. If you only do one LLM learning project this year, make it this one. The understanding you gain from building each component by hand pays dividends every time you debug a model behavior or optimize an inference pipeline.