Definition

Perplexity measures how well an autoregressive language model predicts a sequence of tokens. For each token in a text sample, take the probability the model assigned to that token given the preceding tokens. Average the negative natural log probabilities, then exponentiate the result:

PPL = exp[-(1/N) Σ_i log p(x_i | x_<i)]

Here, N is the number of scored tokens, and p(x_i | x_<i) is the model’s probability for the observed token given its prefix. This is the exponentiated average next-token cross-entropy loss. Lower perplexity means the model assigned more probability to the observed tokens on average.

Simple example

Suppose a two-token test sequence has conditional probabilities of 0.5 and 0.25 under a model. Its perplexity is exp[-(log 0.5 + log 0.25) / 2], or about 2.83. If the model assigned 0.5 to both observed tokens, perplexity would be 2.

The score describes those predictions on that text. It does not say whether either token makes the text useful or true.

Why it matters

Perplexity tracks next-token prediction on held-out text. During training, falling perplexity on the same, consistently scored evaluation set means the geometric mean of the probabilities assigned to its tokens has risen. Keep the dataset and scoring procedure beside the number. Without them, the score is hard to interpret.

One important nuance

Perplexity is measured per token, and token boundaries depend on the tokenizer. Two models can split the same text differently, so their raw perplexity numbers are not necessarily comparable. The evaluation text, preprocessing, context length, and how windows are scored also affect the result. Keep those conditions consistent before treating a lower number as an improvement. Even then, lower perplexity does not establish better factual accuracy or performance on an application task.