Definition
Decode is the output-generation phase of autoregressive language model inference. In a typical decoder-only transformer, prefill processes the prompt and produces the scores used to select the first output token. Once that token is selected, ordinary decoding repeatedly scores the next-token candidates conditioned on the preceding sequence, selects a token, and extends the output. The loop ends when a stopping condition is met.
For a decoder-only transformer, the runtime typically reuses the prompt’s KV cache and extends it with state for generated tokens. It does not have to recompute the entire prompt at every step.
Simple example
A request asks for an explanation of a failed deployment. Prefill supplies the scores for the first output token. If the response continues for 80 more tokens, ordinary decoding takes 80 further steps for that sequence. Each step uses the prompt and the output tokens chosen so far. The model cannot generate those 80 positions independently in one ordinary step.
Why it matters
Every additional output token adds work. A request with a short prompt and a long answer may start quickly but take much longer to finish. Time per Output Token measures the pace after output begins. Time to First Token measures the wait until the first token arrives.
KV-cache reuse reduces repeated computation, but the cache occupies memory and generally grows as more tokens are retained. Output limits cap generation length, helping constrain decode work and KV-cache memory usage.
One important nuance
“One token at a time” describes ordinary autoregressive decoding for each sequence, not how many requests a server can run together. Continuous batching can combine decode steps from several requests. Speculative decoding can also verify proposed tokens in parallel and sometimes advance one sequence by more than one token per target-model pass. In both cases, later accepted tokens still depend on the earlier accepted prefix.