Large Language Models
Scaling Laws
Predicting capability from parameters, data, and compute.
You can jump straight in, but this star assumes Pretraining. Not recommended as a first stop.
Four layers of depth
Each layer ends with a quiz. Finish layer 4 and you own this concept.
- L1IntuitionWhy loss curves look boringly predictable, and how a handful of power-law exponents rewrote how the entire industry spends its compute budget.6m
40 XP - L2MechanicsThe machinery behind the curve: the parametric loss form, why Chinchilla's fit differs from Kaplan's, and how the compute-optimal allocation falls out of it.9m
70 XP - L3CodeFit a power law to real loss-vs-compute data with scipy, then reproduce the Chinchilla-style compute-optimal N/D split numerically.10m
110 XP - L4FoundationsThe full parametric loss surface, the Lagrangian derivation of the compute-optimal split, and why loss extrapolation has hard limits.10m
180 XP
33 stars in the atlas.