Definition
Dynamic batching is a serving technique that combines separate inference requests into batches at serving time. The scheduler draws from requests already waiting and may hold them briefly so more can arrive. A batch can be dispatched when it reaches a configured size, another dispatch condition is met, or the batching wait expires. Execution still depends on available capacity.
The callers submit individual requests and receive individual results. The serving runtime handles the grouping.
Simple example
Suppose a classification service can process up to eight inputs in one model execution. A request arrives, followed by five more during a configured 2 ms wait. The scheduler can group those six inputs rather than waiting for all eight. If only one arrives, the 2 ms limit can end batching with a single input. That request may still wait for execution capacity.
The 2 ms is an example of a scheduling limit, not a recommended setting. A busy service may already have enough queued work to form a batch without waiting.
Why it matters
Grouping requests can make better use of an accelerator and increase throughput when the model and runtime benefit from batched execution. Waiting to fill a batch adds queue time to some requests, though. For an interactive service, that delay may outweigh the throughput gain.
Tune the batch size and wait limit against traffic that resembles production. Measure both completed requests per second and request latency, especially the slower percentiles. At low traffic, the batcher may send mostly single requests. Under heavier traffic, queueing can still grow even with full batches.
One important nuance
Dynamic batching decides which arriving requests share an execution. Continuous batching can also change the group between generation steps, admitting new sequences as others finish. A runtime may use both ideas, but forming a batch from a short arrival window alone does not imply iteration-level scheduling.