Definition
Online inference runs a trained model in response to a request as it arrives. A person or another service expects a result promptly, so the serving path has a latency target set by the use case. The serving system might return a classification in one response or stream generated text over several seconds. “Online” describes when the work is triggered and how quickly the caller expects a result. It does not mean the model learns from the request.
Simple example
A support agent opens a ticket and asks for a draft reply. The application sends the ticket and instructions to a model, then shows the generated text as it arrives. The agent can read the opening lines while generation continues. If the request spends 20 seconds in a queue, fast token generation afterward will not make the interaction feel responsive.
Why it matters
An online request can spend time waiting for capacity, processing its input, generating output, and crossing the network. All of that time counts from the caller’s perspective. Measure latency under realistic concurrency. An isolated model call may miss the queueing caused by concurrent requests.
For streamed text, measure Time to First Token from when the caller sends the request until it receives the first generated token. Measure completion latency from that same start until the full response arrives. Streaming can make the wait easier to tolerate, but it does not remove queueing or shorten the work needed to complete the answer.
One important nuance
An online service can batch requests from different callers internally. Each caller still waits for its own result. Streaming is optional: a classifier may return one score, while a language model may emit tokens progressively. When batching is used, serving policies balance throughput against per-request latency targets.