Definition

Sampling chooses an output token by drawing from a probability distribution over possible candidates. Tokens with higher probabilities are more likely to be selected, but less likely tokens can still be chosen.

In an autoregressive language model, the distribution depends on both the input and the tokens generated up to that point. After the decoder samples a token, it appends that token to the sequence. The model then generates scores for the next position.

Simple example

Suppose the prefix is The status is, and a simplified distribution allows only these candidate tokens:

Next tokenProbability
ready60%
pending30%
blocked10%

Assume, for this example, that each entry represents a single token. These probabilities are hypothetical.

Over many independent draws from this same distribution, ready would appear roughly 60% of the time. A single draw could still select blocked. Ten draws will not necessarily yield exactly six, three, and one of each candidate.

After a token is selected, the next probability distribution depends on that choice. The decoder does not reuse this table for future tokens in the response.

Why it matters

Sampling lets repeated requests with the same prompt generate different continuations. Use it when your application needs alternative phrasings or multiple candidate answers. For evaluation, run repeated tests instead of judging a configuration based on a single response.

Temperature adjusts the relative probabilities used during sampling. Top-k and top-p limit the eligible candidates, then normalize their probabilities before drawing. These controls shape sampling behavior.

One important nuance

A probabilistic model does not require random token selection. Greedy decoding always picks the highest-probability token at each step instead of drawing from the distribution.

Neither method guarantees factual accuracy. A likely continuation can still contain incorrect information, so do not treat sampling settings as a substitute for validation.