Definition

A logit is a raw score a language model assigns to a candidate next token. At each generation step, the model produces one for every token in its vocabulary. Logits can be negative, and they do not have to add up to one.

Softmax turns those scores into a probability distribution. It exponentiates each logit and divides by the sum of all exponentiated logits. A higher logit gives a token a higher probability relative to the other candidates at that step.

Simple example

Imagine a vocabulary with only three tokens. The model gives them logits of 2, 1, and 0. Softmax turns those into probabilities of about 0.665, 0.245, and 0.090.

The first token has the highest score, but its probability is about 66.5 percent, not 2 percent. With a larger vocabulary, the probabilities also depend on the scores of all other tokens included in the softmax calculation.

Why it matters

During temperature-controlled sampling, the inference runtime typically scales the model’s logits before computing sampling probabilities. If an inference API exposes token scores, check whether it returns logits, probabilities, or log probabilities before applying a threshold. A logit of 0.8 does not mean an 80 percent probability.

One important nuance

The gaps between logits matter more than their absolute values. Adding the same number to every logit leaves the softmax probabilities unchanged: 2, 1, 0 and 102, 101, 100 produce the same distribution. A lone logit therefore tells you little about a token’s probability without the scores of the other candidates.