Definition

Prefill is the phase of autoregressive language model inference that processes the input tokens before the model continues with generated tokens. In a decoder-only transformer with causal self-attention, the model computes representations for the known prompt positions. The serving runtime typically caches their attention keys and values, which later decoding steps can reuse.

The prompt is already available, so the model can process its positions in parallel while each position attends to itself and preceding positions. The final prompt position supplies the scores used to select the first output token. After that, the model generates further tokens incrementally.

Simple example

An application sends a 2,000-token prompt containing an incident report and an instruction to summarize it. During prefill, the model processes those input tokens and builds the prompt’s KV cache. The serving runtime uses the resulting scores to select the first output token. To produce the rest of the summary, each decode step uses the cached prompt state and extends it with generated tokens.

With no cached prefix, the runtime processes the whole prompt before the summary can start. Later decode steps reuse that state rather than recomputing it.

Why it matters

Prefill contributes to the wait before the first output token. Longer prompts generally require more prefill work, so adding retrieved documents or long tool descriptions can delay the start of a response even when the answer is short. That makes input length worth measuring alongside Time to First Token.

Separating prefill time from queueing and decoding time helps identify which part of an inference request is slow. Client-observed Time to First Token includes more than prefill, such as waiting for capacity and delivering the first token to the client.

One important nuance

Prefill is a phase of work, not necessarily one uninterrupted model pass. A serving runtime may split a long prompt into chunks or reuse cached state for a matching prefix. A 2,000-token prompt can therefore require different amounts of fresh prefill work across requests.