Definition

Entropy measures how spread out the probabilities in a distribution are. In next-token prediction, the distribution contains the model’s probabilities for the possible tokens, given the current context. Concentrating probability on one token gives low entropy. Spreading it across many tokens gives higher entropy.

For token probabilities p(t), the calculation is H(p) = -Σ p(t) log₂ p(t). Base 2 reports entropy in bits. Natural logarithms report it in nats. A token with probability zero contributes zero by convention.

Simple example

Suppose a toy model has only two possible next tokens: yes and no. If it assigns each a probability of 0.5, the entropy is 1 bit. If it assigns 0.9 to yes and 0.1 to no, the entropy falls to about 0.47 bits. The second distribution is more concentrated, even though both still allow either token.

An unconstrained language model assigns probabilities across its output vocabulary. Its next-token distribution changes as the context grows.

Why it matters

Two contexts can favor the same next token yet have different entropy because the remaining probability is distributed differently. Entropy summarizes all token probabilities, but it cannot tell you the chance of sampling a particular alternative.

When comparing entropy across contexts or decoding settings, measure it at the same stage. Temperature and token filtering can reshape the distribution. Mixing values calculated before those controls with values calculated after them can mislead.

One important nuance

Low entropy does not mean the likely token is correct. A model can assign nearly all its probability to the wrong token. Entropy measures concentration, not whether predicted probabilities match observed outcomes or whether the resulting text is factually sound.