Pretraining

Layer 4 · Foundations

Pretraining

Cross-entropy and perplexity derived properly, the Chinchilla compute-optimal arithmetic, and where the FLOPs actually go on real hardware.

15 min read180 XP

Cross-entropy as maximum likelihood

Pretraining is maximum likelihood estimation over sequences. For a corpus of tokens , the model defines a probability for the whole sequence via the chain rule, and training maximizes it — equivalently, minimizes its negative log: