Definition

Cross-entropy scores a model’s predicted probabilities against a target distribution. For target probabilities q and model probabilities p, it is H(q, p) = -Σ q(x) log p(x). Each outcome’s contribution depends on its target probability, so a low model probability matters more for outcomes with high target probability. The logarithm’s base determines the unit: natural logs give nats, while base 2 gives bits.

In next-token training, the target is usually one observed token. Its target distribution assigns that token probability 1 and every other token 0. The loss for that position then reduces to -log p(target token). Training commonly averages this loss across selected token positions. The same calculation on held-out text can measure prediction loss without updating the model.

Simple example

Suppose the observed next token is yes. A model assigns yes probability 0.8 and no probability 0.2. With natural logs, the cross-entropy loss is -ln(0.8), about 0.223 nats. If the model assigns only 0.2 to yes, the loss rises to -ln(0.2), about 1.609 nats. The other token matters because its probability leaves less for the observed one.

Why it matters

During training, the optimizer uses this loss to adjust model weights. On a fixed evaluation set, a lower average loss means the model’s geometric mean probability for the observed tokens is higher. Keep the data, tokenization, and scoring procedure consistent across runs. That comparison measures prediction, not whether the text is true or useful.

Perplexity is e raised to average next-token cross-entropy measured in nats, or 2 raised to the average measured in bits.

One important nuance

Cross-entropy is not always zero when predictions match the target distribution. If a target assigns 0.5 to each of two outcomes and the model does the same, its cross-entropy is 1 bit. That is the target distribution’s own entropy, and no prediction can do better for that target.