Sebastian Raschka Β· Learn by building
Super Intelligence (SI)
From Scratch
I like to understand a model by implementing it. This page collects my from-scratch material, from neural networks and LLMs to reasoning methods and coding agents.
Start with the books
All my books
01 / LLM foundations
Build a Large Language Model (From Scratch)
Build a GPT-style LLM in Python and PyTorch, from tokenization and attention to pretraining and finetuning.
Book & resources
02 / Reasoning models
Build a Reasoning Model (From Scratch)
Start with a pretrained LLM. Implement evaluation, inference-time scaling, reinforcement learning, and distillation.
Book & resourcesWhere to begin? Start with the LLM book for the full implementation path. The reasoning book can also be read independently. Both have freely available code and exercises.
-
Free video course
LLMs from scratch video course
Code along with the LLM book, chapter by chapter.
Code -
Talk
Implementing LLM architectures from scratch
How to implement open-weight model architectures and compare them with reference code.
-
Free video course
Reasoning models from scratch
Code along with the reasoning book, from loading a base model to implementing reasoning methods.
Essentials
-
Chapter code
Working with text data
Turn text into tokens, embeddings, and input-target pairs for training.
-
Article + code
A BPE tokenizer from scratch
Implement byte pair encoding and work through the tokenizer's training and merge rules.
Code -
Article + code
Self-attention from scratch
A step-by-step implementation of scaled dot-product attention in PyTorch.
-
Article + code
Multi-head, causal, and cross-attention
Extend the attention mechanism to multiple heads, causal masking, and cross-attention.
-
Chapter code
Implementing a GPT model
Combine the individual building blocks (attention, feed-forward networks, and so on) into a GPT model.
Advanced concepts
-
Article + code
The KV cache from scratch
Cache keys and values to avoid recomputing them at every text-generation step.
Code -
Code comparison
Efficient multi-head attention
Compare several PyTorch implementations of the same attention mechanism.
-
Guide + code
Grouped-query attention (GQA)
Share key and value heads across groups of query heads.
Code -
Guide + code
Multi-head latent attention (MLA)
Compress key and value information into a smaller latent representation.
Code -
Guide + code
Sliding-window attention
Restrict attention to a local window of preceding tokens.
Code -
Guide + code
Mixture of experts (MoE)
Route each token to a subset of expert feed-forward networks.
Code -
Guide + code
Gated DeltaNet
Explore recurrent state updates as an alternative to storing a growing KV cache.
Code -
Guide + code
DeepSeek Sparse Attention
Use an indexer to select a subset of tokens for attention.
Code -
Code
Cross-layer KV sharing
Reuse key and value states across transformer layers.
-
Concept guide
RMSNorm
Normalize activations using their root mean square.
-
Concept guide
Rotary positional embeddings (RoPE)
Add positional information to the token embeddings via rotation.
-
Concept guide
SiLU and gated feed-forward layers
Non-linear activation functions used in the feed-forward portions.
-
Concept guide
Query-key normalization
Normalize queries and keys before computing attention scores.
-
Concept guide
No positional embeddings (NoPE)
How causal transformers can operate without explicit positional embeddings.
-
Concept guide
Gated attention
Add learned gates to control the attention output.
-
Model code
Llama 3 and 3.2
Convert the GPT implementation to Llama, including GQA, RoPE, and RMSNorm.
-
Article + model code
Qwen3
Implement the dense and mixture-of-experts Qwen3 architectures.
Code -
Model code
Gemma 3
Implement the text model with its local and global attention layers.
-
Model code
Olmo 3
Build the 7B and 32B architectures and load their weights.
-
Model code
Tiny Aya
A from-scratch implementation of the 3.35B multilingual model.
-
Model code
Qwen3.5
Work through a hybrid model that combines attention and Gated DeltaNet.
-
Model code
Gemma 4 E2B and E4B
Implement the text backbones of the Gemma 4 edge models.
-
Code experiments
Pretraining with other LLM architectures
Use models such as Llama and Qwen as replacements in the Chapter 5 training code.
-
Chapter code
Pretraining on unlabeled text
Implement the loss function, training loop, and text-generation sampling methods.
-
Code project
Pretraining on Project Gutenberg
Prepare a larger text corpus and use it to pretrain the GPT model.
-
Code
Learning-rate schedules and training refinements
Add warmup, cosine decay, and gradient clipping to the training loop.
-
Code experiments
Faster LLM training in PyTorch
Compare practical changes that improve the training code's performance.
-
Code
Memory-efficient weight loading
Reduce peak memory use when loading pretrained model weights.
-
Code experiments
Hyperparameter tuning for pretraining
Run experiments with different training settings for the GPT model.
-
Article + code
A GPT-style text classifier from scratch
Finetune a pretrained language model for text classification.
Code -
Chapter code
Instruction finetuning
Prepare instruction-response examples and train an LLM to follow instructions.
-
Article + code
LoRA and DoRA from scratch
Implement low-rank adaptation and its weight-decomposed variant in PyTorch.
-
Tutorial + code
Direct preference optimization (DPO)
Use preferred and rejected responses to implement preference finetuning.
-
Code
Generating instruction datasets
Generate and improve the instruction data used for finetuning.
-
End-to-end project
An AI text detector from scratch
Build a dataset, train a classifier, deploy it locally, and explore RLVR.
-
Book chapter preview
A first look at reasoning from scratch
The conceptual starting point for the reasoning book and its implementation path.
-
Chapter code
Generating text with a pretrained LLM
Load a Qwen3 model and generate text before adding reasoning methods.
-
Article + code
Four approaches to LLM evaluation
Implement multiple-choice benchmarks, verifiers, leaderboard comparisons, and LLM judges.
Code -
Chapter code
Evaluating reasoning with verifiers
Parse model responses and check their answers on math problems.
-
Chapter code
Inference-time scaling
Generate multiple candidate answers and compare selection methods.
-
Chapter code
Self-refinement
Give the model opportunities to revise its own responses.
-
Chapter code
Reinforcement learning with verifiable rewards
Implement GRPO to train a reasoning model using rewards from an answer verifier.
-
Chapter code + experiments
Improving GRPO
Explore changes to policy optimization and compare advanced GRPO training scripts.
-
Chapter code
Distilling a reasoning model
Generate teacher responses and use them to train a smaller reasoning model.
Complete list of reasoning model code, exercises, and bonus materials
-
Article + code
Components of a coding agent
Understand the building blocks of a coding agent harness.
Code -
Practical guide
Using local coding agents
How to set up an open-weight LLM in a local coding harness, on a local machine.
-
Code project
Building a chat interface
Add a chat interface to interact with the reasoning model.
-
Tutorial + code
Neural networks and gradient descent
Work through artificial neurons, single-layer networks, and their training rules.
-
Notebook
Backpropagation from scratch
Implement a multilayer perceptron and compute its gradients explicitly.
-
Tutorial + code
Principal component analysis
Build PCA from the covariance matrix, eigenvectors, and a projection onto principal components.
-
Tutorial + code
Linear discriminant analysis
Implement supervised dimensionality reduction using within-class and between-class scatter.
-
Tutorial + code
Kernel PCA
Extend PCA to nonlinear data using an RBF kernel.
-
Tutorial
Naive Bayes and text classification
Derive a probabilistic text classifier from Bayes' theorem.
-
Notebook collection
Deep learning model zoo
Standalone implementations of neural networks, CNNs, recurrent models, autoencoders, and more.
-
FAQ
Why implement algorithms from scratch?
When implementing a method helps with understanding, experimentation, and checking assumptions.