Table of Contents
When a model slows down, the AI service must keep admitted work within usable capacity. Accept less work or add more capacity. Otherwise, unfinished work piles up until the rest of the product is affected. Limit any waiting to the period when the result is still valuable. Decide whether to reject a request before starting costly downstream operations.
Backpressure occurs when downstream capacity constraints cause upstream producers to slow down, wait, or stop admitting work. A worker may slow its reads, or a producer may wait for available space. An API can signal overload by rejecting requests. Load shedding becomes effective as a backpressure mechanism only when the producer responds by reducing its load. If callers immediately retry, the server continues to shed load while producers keep sending requests.
Slow calls turn into accumulated work
Consider a support-reply feature: it loads a ticket, retrieves policy documents, generates a draft, and validates the result. A batch worker may use the same model deployment to summarize older tickets.
If the model slows down while incoming traffic remains steady, fewer operations complete. Requests start to accumulate. Some callers may time out and retry, even as their original work continues to run.
A queue can hold unfinished work until a worker is available to process it, subject to capacity, expiration, and drop policies. If capacity becomes available later, this approach buys time during a burst.
For example, if the system receives eight jobs per second but completes only five, the backlog grows by three per second and reaches 180 jobs after one minute. At five completions per second, a new arrival would face about 36 seconds of queued processing.
If arrivals drop to four per second, the backlog clears at a rate of one job per second. It would take about three minutes to eliminate the backlog. Processing times vary, but this illustrates how a short spike can cause a long delay.
The Amazon Builders’ Library explains how timeouts and retries can worsen overload. It also recommends limiting queue age, so workers can discard requests that are no longer useful, rather than processing them after the caller has already given up.
Give each control a specific job
Runtime budgets bound one execution, while shared capacity controls bound the work competing with that execution.
| Control | Decision it owns | Limitation |
|---|---|---|
| Execution budget | How much time, model work, and spend may this operation use? | Many individually bounded operations can still overload a service |
| Concurrency limit | How many operations may hold this resource at once? | Fast operations can still exceed a time-based quota |
| Rate limit | How much work may start over an interval? | Slow operations can accumulate even at an allowed arrival rate |
| Bounded queue | How much waiting work may the system retain? | Buffered work still needs processing capacity and a useful lifetime |
| Admission control | Should the application accept this work now? | The decision needs a defined overload outcome |
Admission control determines whether to accept work based on capacity and policy signals. A single concurrency threshold can often suffice. Load shedding deliberately rejects or removes work to protect capacity. Backpressure completes the loop when producers respond by reducing their load.
Request count is a rough measure of load. A short classification and a long summary can consume very different amounts of model time and tokens.
Separate workloads where cost differences are significant. A request-counting token bucket does not track an LLM provider’s token quota. Admission to provider calls may require estimated token reservations and request-rate limits. The application can set its own concurrency limit to bound active work. This does not have to match the provider’s enforced quota.
Reject at the HTTP boundary before starting costly work
Consider the support-reply feature from earlier. Suppose it exposes an HTTP endpoint that generates a draft for a ticket. For this endpoint, set a specific concurrency limit and use no middleware queue. Reject excess requests before retrieval or generation begins, and document this behavior clearly.
These examples target .NET 10 with the ASP.NET Core framework. Replace the simulated delay with your application service call and keep its cancellation token.
using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRateLimiter(options =>
{
options.AddConcurrencyLimiter("interactive-ai", limiter =>
{
limiter.PermitLimit = 8; // Illustrative, per process.
limiter.QueueLimit = 0;
});
options.OnRejected = async (context, cancellationToken) =>
{
await Results.Problem(
statusCode: StatusCodes.Status503ServiceUnavailable,
title: "AI capacity is busy",
detail: "This request was not accepted. Try again later.")
.ExecuteAsync(context.HttpContext);
};
});
var app = builder.Build();
app.UseRouting();
app.UseRateLimiter();
app.MapPost("/tickets/{id:guid}/draft", async (
Guid id, CancellationToken cancellationToken) =>
{
await Task.Delay(TimeSpan.FromSeconds(2), cancellationToken);
return Results.Ok(new { ticketId = id, draft = "Example reply" });
}).RequireRateLimiting("interactive-ai");
app.Run();
The middleware holds a permit for the duration of the request pipeline. With QueueLimit = 0, excess requests are rejected immediately. This example uses 503 for shared capacity shortages. If a caller exceeds only its own allowance, a 429 Too Many Requests may be more appropriate. HTTP specifications do not require this distinction. Document your choice in the API contract.
Include a Retry-After header only when you have a meaningful retry time. A concurrency limiter cannot predict when active work will complete, so it cannot provide an estimate. Do not guess a recovery time based on permit count. Clients should use bounded retries and backoff, even without this header.
This policy pool exists within a single process. Eight permits on each of four replicas allow up to 32 concurrent requests across all replicas. Per-tenant limits help prevent one tenant from monopolizing capacity. An aggregate limit protects overall capacity, since multiple tenants staying within their own limits can still exceed the total.
Put a second gate where callers share the model
The endpoint limit applies to interactive requests. A batch worker may bypass this limit, and a single admitted request could initiate several model calls in parallel.
Place a gate around each model attempt. The following class limits active attempts for a single capacity pool within one process, so interactive services and workers in that process must share the same instance.
If the HTTP service and worker run in separate processes, each will have its own gate. For example, two gates with four permits each allow eight concurrent attempts. This class does not enforce a combined provider limit across processes:
using System.Threading.RateLimiting;
public sealed class ModelCapacity : IDisposable
{
private readonly ConcurrencyLimiter limiter;
public ModelCapacity(int maxConcurrentAttempts)
{
limiter = new ConcurrencyLimiter(new ConcurrencyLimiterOptions
{
PermitLimit = maxConcurrentAttempts,
QueueLimit = 0
});
}
public async Task<T> RunAsync<T>(
Func<CancellationToken, Task<T>> attempt,
CancellationToken executionToken)
{
executionToken.ThrowIfCancellationRequested();
using var lease = limiter.AttemptAcquire(1);
if (!lease.IsAcquired)
throw new ModelCapacityUnavailableException();
executionToken.ThrowIfCancellationRequested();
return await attempt(executionToken);
}
public void Dispose() => limiter.Dispose();
}
public sealed class ModelCapacityUnavailableException : Exception
{
public ModelCapacityUnavailableException()
: base("No model capacity is available.") { }
}
Register one shared instance through dependency injection:
builder.Services.AddSingleton<ModelCapacity>(
_ => new ModelCapacity(maxConcurrentAttempts: 4));
Four is illustrative. Creating a gate per request would give each request its own allowance. Register the shared instance as a singleton and let the container own its disposal.
AttemptAcquire returns immediately. If a lease is not acquired, the delegate does not run. The using scope holds the acquired permit until the attempt completes, throws an exception, or is cooperatively canceled.
Call RunAsync within the provider adapter for each awaited attempt, passing the execution-deadline token. For interactive calls, map rejection to an unavailable result. For worker jobs, map it to a bounded retry under the job policy. Avoid tight retry loops and keep retry policies explicit.
For streaming, hold the lease through enumeration and disposal. Returning an IAsyncEnumerable<T> from this delegate releases the permit before the caller consumes it. This helper requires a task that represents the complete operation.
Cancellation asks work to stop. It does not prove a remote provider stopped processing or billing. Keep the local lease until the underlying task actually ends, even if a separate timeout race completed earlier.
A production gate may also enforce request rate, token reservations, and throttle cooldowns. Keep independent capacity pools separate. Across processes or replicas, use distributed limiting or controlled aggregate worker dispatch, because more callers cannot enlarge a fixed downstream quota.
Decide how long waiting remains useful
An interactive request can sometimes tolerate a short wait. Count that wait against its original deadline. If the application starts its runtime budget after admission, time in a limiter queue can disappear from the budget while the caller still experiences it.
Prefer one deliberate waiting point per execution path. Stacked HTTP, service, and provider waits can consume the deadline before generation begins. Measure their combined duration.
For durable jobs, bound backlog size and age. Before dispatch, check the deadline, input version, and cancellation state. Recheck current authorization if the job’s execution policy requires it. Record an explicit terminal outcome for expired or superseded jobs.
Monitor the age of the oldest pending job alongside queue depth. A hundred classifications may clear much faster than a hundred summaries. If estimated completion no longer fits the product’s promise, stop admitting jobs or offer a later completion contract the caller can accept.
Make the accepted work durable before returning 202 Accepted. An HTTP limiter queue provides no durability. When an AI request should become a background job covers ownership and lifecycle.
For disposable or recoverable local work, use a bounded Channel<T>. Choose its waiting, rejection, or drop policy. Await writes and bound processing concurrency. Also bound or reject producers where necessary so a full channel does not turn into an unbounded set of pending writes.
Make retries and recovery obey the same limits
A throttle affecting a shared provider pool should change admission for every caller using it. Otherwise, other HTTP requests and batch jobs keep sending into the same shortage.
Normalize the response in the adapter and update admission or cooldown state for the affected pool. Treat provider throttling as a shared capacity signal covers this boundary. A caller-specific limit should leave unrelated traffic alone.
Release the attempt’s concurrency permit before a retry delay, then re-enter admission for the next attempt. The delay and repeated work stay inside the execution budget. Check SDK and HTTP resilience settings too. Internal SDK retries can bypass per-attempt rate, token, cooldown, or attempt-count checks between physical calls. The outer lease still bounds the logical operation’s concurrency. The overall execution deadline still applies when its cancellation token reaches the SDK and the SDK honors it.
On recovery, resume dispatch gradually under the shared limit. Jitter spreads retries. If you release all delayed calls and queued jobs together, you can recreate the spike.
Reserve interactive capacity when batch work must not crowd it out. Separate queues need scheduling or allocations within the combined allowance. Keep status polling and cancellation usable under saturation so callers can manage accepted jobs.
Test beyond the point where rejection starts
Test beyond expected traffic. Slow the provider and exceed the admission limit. Accepted requests should still finish within their target while excess work receives a prompt refusal.
Record offered load, accepted throughput, and outcomes. Latency among successful requests can look healthy while most callers receive rejections.
Track active attempts per pool, stage wait times, oldest-job age, expired work, throttles, and attempts per completed operation. Use bounded labels such as workload class and configured pool. Keep prompts and arbitrary user identifiers out of metric dimensions.
Exercise the failure paths that change capacity:
- A model attempt fails or cancels. Its local permit becomes available after the attempt exits.
- Interactive traffic and batch jobs compete. Their combined dispatch stays within the intended pool limit.
- A queued job expires before dispatch. It never starts a model call and exposes an expired outcome.
- Replicas scale out. The aggregate provider allowance still holds.
- A provider recovers with a backlog waiting. Dispatch ramps up without an unrestricted replay burst.
- A caller receives an overload response and backs off. Its submission rate drops instead of immediately replacing every rejected request.
Choose limits using realistic input and output sizes, streaming durations, and model-call fan-out. A setting tested only with tiny prompts may fail on users’ actual documents.
When to use backpressure and when not to queue
Use shared controls where AI workloads compete for capacity, especially when interactive requests share a provider with batch work or evaluations. Add waiting only where the product can tolerate it.
Queue work that remains useful later, with defined acceptance, expiry, and processing policies. A durable queue can absorb a burst if later spare capacity drains it within that promise.
Do not queue an interactive answer past its useful deadline, or accept a batch backlog that grows indefinitely under normal demand. Reduce admitted work, obtain more usable capacity, or change the product’s completion promise. A larger buffer will only postpone the same decision.
Pick one endpoint and its busiest background producer. Bound their shared capacity and define what callers see when it runs out. Slow the dependency, then check that accepted work finishes and producers respond to refusals by reducing load.
Related reading
- AI Systems Need Runtime Budgets
- When an AI Request Should Become a Background Job
- Treat provider throttling as a shared capacity signal
Sources
- Amazon Builders’ Library: Using load shedding to avoid overload
- Google SRE: Handling overload
- Microsoft Learn: Rate limiting middleware in ASP.NET Core
- Microsoft Learn: Rate limiter samples and Retry-After limitations
- Microsoft Learn: RateLimiter API and immediate permit acquisition
- RFC 6585, section 4: 429 Too Many Requests
- RFC 9110, sections 10.2.3 and 15.6.4: Retry-After and 503 Service Unavailable