Atlas

Large Language Models

Training at Scale

Splitting one model across thousands of accelerators.

You can jump straight in, but this star assumes Pretraining. Not recommended as a first stop.

Four layers of depth

Each layer ends with a quiz. Finish layer 4 and you own this concept.

  1. L1IntuitionWhy a 70B-parameter model can't be trained on one GPU, and the four different ways of cutting the job into pieces that fit.7m
    40 XP
  2. L2MechanicsZeRO's three stages, tensor and pipeline parallelism mechanically, collectives, and how 3D parallelism combines all of it.11m
    70 XP
  3. L3CodeA real FSDP training loop with activation checkpointing and mixed precision, plus a minimal from-scratch simulation of gradient all-reduce.13m
    110 XP
  4. L4FoundationsMemory accounting across ZeRO stages, all-reduce communication volume, pipeline bubble fraction, and the real interconnect numbers behind all of it.12m
    180 XP

33 stars in the atlas.