Large Language Models
Training at Scale
Splitting one model across thousands of accelerators.
You can jump straight in, but this star assumes Pretraining. Not recommended as a first stop.
Four layers of depth
Each layer ends with a quiz. Finish layer 4 and you own this concept.
- L1IntuitionWhy a 70B-parameter model can't be trained on one GPU, and the four different ways of cutting the job into pieces that fit.7m
40 XP - L2MechanicsZeRO's three stages, tensor and pipeline parallelism mechanically, collectives, and how 3D parallelism combines all of it.11m
70 XP - L3CodeA real FSDP training loop with activation checkpointing and mixed precision, plus a minimal from-scratch simulation of gradient all-reduce.13m
110 XP - L4FoundationsMemory accounting across ZeRO stages, all-reduce communication volume, pipeline bubble fraction, and the real interconnect numbers behind all of it.12m
180 XP
33 stars in the atlas.