Cross-entropy as maximum likelihood
Pretraining is maximum likelihood estimation over sequences. For a corpus of tokens , the model defines a probability for the whole sequence via the chain rule, and training maximizes it — equivalently, minimizes its negative log: