An AI feature does not have to return an answer during an outage.

If it cannot meet its normal contract, a clear unavailable state may be the best response. Other features can fall back to ordinary application code, accept the work for later, keep an independent capability available, or show an older result with a visible warning.

Decide that behavior as part of the product contract. A catch block can detect a timeout, but it cannot decide what the user should still be able to do.

Degrade the user flow, not the exception

Consider a support workspace that can retrieve policy documents and prepare a reply for a ticket. The model call is one part of a larger user flow:

open the ticket
    -> review customer history
    -> retrieve approved policy
    -> prepare an AI draft
    -> edit and save the reply

If the model is unavailable, the human support agent should still be able to open the ticket, inspect its history, and write a reply manually, provided those capabilities do not depend on the model. Disabling the entire workspace because one optional capability failed would turn a dependency incident into a larger product outage.

Different dependency failures call for different behavior across user flows. A live suggestion can disappear without much harm. A long-running document analysis may be worth queueing. An answer that must be grounded in current policy should stop when the application cannot obtain that policy evidence, even if the model still responds.

Start with what the user needs to accomplish, then decide how the flow should behave during each failure.

For each important flow, define:

  • what remains useful without the failed dependency
  • which result must no longer be produced
  • whether a safe substitute exists
  • whether the work still has value later
  • what the user needs to know about the changed behavior

Write those decisions down as the degradation contract for the flow.

Five useful degradation modes

Start with five possible product behaviors, then add failover machinery where the product needs it.

ModeUse it whenExample
Explicit unavailable stateNo safe or useful substitute existsDisable AI reply generation while leaving the ticket editor available
Deterministic non-AI pathRules, templates, or ordinary application code can still complete a narrower jobInsert an approved acknowledgement template instead of generating a reply
Delayed processingThe result remains useful after the current request endsAccept a document summary job and expose its status
Reduced functionalityOne independent part of the feature still has valueShow retrieved policy documents without generating a synthesis
Disclosed stale resultA previous result remains relevant within a defined age and input versionShow an earlier draft for the same ticket revision and label when it was produced

These modes can coexist in one product. The workspace might display an earlier draft, offer an approved template, and accept a new drafting job at the same time. Give each operation a clear outcome and show which workspace capabilities are currently available.

Return an explicit unavailable state

If the feature cannot produce a result that meets its normal contract, say so. Keep the unaffected application available and make the boundary visible. “AI reply drafting is temporarily unavailable” is more useful than a spinner that eventually disappears or a generic error at the top of the whole page.

Do not replace a grounded answer with an ungrounded one just to avoid showing an error. If the application cannot obtain the required policy evidence from an approved source, it cannot prepare that reply. A failed retriever may leave another approved source available. A responding model does not make missing evidence safe to ignore.

At an HTTP boundary, a temporary inability to handle the operation may map to 503 Service Unavailable. Send Retry-After only when the application can suggest a meaningful interval. An arbitrary fixed interval can make clients retry together. Clients can use jitter to spread their attempts. A per-client rate limit may call for 429 Too Many Requests. Invalid caller credentials call for 401 Unauthorized with a WWW-Authenticate challenge (provided those capabilities do not depend on the model). Valid credentials without sufficient permission generally call for 403 Forbidden. Choose the status from the failure and the API contract, not from the fact that generation did not finish.

Do not blindly forward an upstream provider’s status code. Map dependency failures according to your own API contract so callers can distinguish their rate limits from temporary service-capacity problems.

Use a deterministic non-AI path

A non-AI path works when the product already has a narrower, predictable way to help.

For the support example, the application might offer an approved acknowledgement template. Its text and selection rules can be reviewed and tested. The template does less than a generated draft, and it does not need to imitate one.

Do not present a canned sentence as if the assistant wrote a case-specific answer. Label the alternative for what it is and let the user decide whether it helps.

Deterministic paths are especially useful when AI improves speed but does not own the underlying job. They are much harder to add later if the user interface and application service assume that every successful flow must contain model output.

Delay work that still matters later

A document analysis, report, or batch classification can remain useful after the interactive request ends. The application can durably accept the work, return an operation identifier, and process it when the dependency recovers. An HTTP API commonly represents this with 202 Accepted and a status resource.

Persist the job and its input identity before returning 202. If dispatch uses a separate queue, write an outbox entry in the same transaction as the job or use another reliable dispatch path. Saving the job and then publishing a message leaves a failure window.

Enforce submission idempotency keys atomically, scoped to the caller and operation and bound to the inputs. A retry with the same key and inputs returns the existing job. Different inputs are a conflict. Message delivery can repeat, so workers need an atomic claim and idempotent result persistence. For external calls, use the other system’s idempotency support where available. Otherwise, define how to handle an uncertain outcome before retrying.

Define how long the job may wait, when it expires, whether the user can cancel it, and what terminal failure looks like. Check the ticket revision, policy version, and required permissions before execution and again before making the result available. If required permissions were revoked during generation, do not expose the result. If relevant inputs changed, reject the draft or apply the product’s stale-result policy. Guard the result write with an expected-version condition where the data shares a store. Synchronous generation has the same race.

Authorize status polling, cancellation, and result retrieval against current permissions. A job ID is not proof of access. Revoked ticket access must also block a completed result.

RFC 9110 says 202 means processing was accepted, not that it will eventually succeed. An in-memory task can disappear during a restart.

Do not queue every failed chat message by default. A reply that arrives 20 minutes after the conversation moved on may be worse than an immediate unavailable state.

Keep independent capabilities available

If retrieval works but the model does not, the application may still show retrieved source documents after the usual relevance, freshness, and authorization checks. If required evidence is unavailable, the application may keep unrelated writing assistance available while disabling policy-grounded answers. If optional enrichment fails, the core transaction may continue without it.

Manual editing and policy search are separate workspace capabilities. Keep them available when their own dependencies work. Policy search may itself use embeddings or reranking, so it will not survive every AI outage.

The reply-preparation operation can keep retrieval, generation, validation, and persistence together. Separate unrelated ticket viewing and editing from that operation so a drafting failure does not take down the workspace. The user interface also needs to distinguish “the page failed” from “this capability is temporarily unavailable”.

Reduced mode must preserve the original safety and authorization rules. Dependency failure is not permission to skip access checks, output validation, required audit records, or tenant isolation.

Treat stale results as a separate product mode

For a previous support draft, a time-to-live is not enough. The application should also know that the ticket revision, policy version, tenant, and relevant user context still match. It must recheck current authorization before returning the draft as an accepted previous result. If the ticket changed after the draft was produced, a three-minute-old draft may already be wrong.

Return the time the result was generated and label it as an earlier result. Define a maximum age and the events that invalidate it. Sensitive or consequential answers may not permit stale use at all.

This is stricter than caching a public, slow-changing reference value. Generated content can contain private context and can depend on inputs that are not obvious from the final text. A cache key based only on the user’s question is rarely enough.

Choose the mode from the user promise

For the support flow, a decision table might look like this:

Failed capabilityWhat remains availableChosen behaviorWhy
Model generationTicket view, manual editor, templates, policy searchOffer a template or manual path. Optionally queue a draftThe support agent can still complete the job without pretending a generated reply exists
Required policy evidence unavailableTicket view, manual editor, non-policy toolsDisable grounded reply generationThe application cannot support current policy claims without approved evidence
Durable job acceptance unavailableInteractive pathsDo not offer delayed processingIf the application cannot durably accept and later dispatch the job, it must not return 202.
Previous-draft storeNormal live generationContinue without stale-result modeCached output is an optional convenience, not a critical dependency
Live authorization required but unavailableOperations that can establish authorization with the required guaranteesStop operations that require a live decisionA locally validated token may suffice for some operations, but current permissions or revocation checks may be required for others

If the application cannot establish authorization with the guarantees an operation requires, that operation must stop.

Use these questions when choosing a mode:

  1. Can the user still complete the underlying job without this AI capability?
  2. What harm could a wrong, incomplete, or stale result cause?
  3. Does the result still matter after the current request ends?
  4. Which parts of the flow are genuinely independent?
  5. Can the application explain the reduced behavior without making the user infer it?

If the team cannot answer the second question, default to unavailable until it can.

Put degraded outcomes in the application contract

Callers should not have to parse exception messages to learn whether work was queued, a template was returned, or the feature is unavailable.

Model reply preparation as an application outcome:

public enum ReplyMode
{
    Generated,
    ApprovedTemplate,
    Queued,
    PreviousDraft,
    Unavailable
}

public enum DegradationReason
{
    None,
    ModelUnavailable,
    RetrievalUnavailable,
    RequiredEvidenceUnavailable,
    RequestDeadlineExceeded
}

public sealed record PrepareReplyResult(
    ReplyMode Mode,
    DegradationReason Reason,
    string? Content,
    Guid? JobId,
    DateTimeOffset? GeneratedAt,
    string? InputVersion);

The type is intentionally limited to reply preparation. ReducedFunctionality is a state of the wider workspace, not another kind of reply. Ticket viewing, policy search, and drafting each need their own availability state. RetrievalUnavailable means a required retrieval dependency failed. RequiredEvidenceUnavailable means retrieval ran but no approved source provided enough evidence. An approved template chosen during normal work has Reason.None. One offered because generation failed has Reason.ModelUnavailable.

PreviousDraft means more than “cache hit”. It tells the endpoint and user interface to show the age and reduced-mode notice. Here, InputVersion is a composite identifier for the ticket revision, policy version, tenant, and relevant user context used to produce the draft. Compare those inputs and recheck current authorization before returning it as an accepted result. Queued requires a job identifier and must not contain a pretend final result. Unavailable is a handled outcome, not an unexpected exception.

Production code can use factory methods or separate result types to prevent invalid combinations. The mode should cross the service boundary explicitly.

Map it deliberately at the edge:

  • Generated, ApprovedTemplate, and an accepted PreviousDraft can return a normal response with mode metadata.
  • Queued can return 202 Accepted with a status URL.
  • Unavailable can return a typed application error or 503 for temporary service inability, depending on the API contract. Handle per-client rate limits with 429 and invalid credentials with 401 at the appropriate boundary.

RequestDeadlineExceeded means this request ran out of time. A token limit, provider rate limit, or monetary budget needs a different reason and may require a different response.

The user-facing message should describe what changed and what remains possible. Keep provider names and circuit state out of the UI. Record structured failure information in telemetry, and sanitize exception details so prompts, credentials, or customer data do not leak into logs.

Observe degraded operation separately from failure

A request can be technically successful while the product is degraded. If a template returns with HTTP 200, an availability dashboard that counts only status codes will miss the model outage.

Record enough structured data to distinguish the paths:

  • requested capability and actual reply mode
  • degradation reason and failed dependency category
  • whether work was queued
  • age and input version for a stale result
  • time spent before the system changed modes

Do not copy prompts, ticket content, or generated replies into telemetry to explain the mode. Stable identifiers and outcome categories should be enough for operational analysis.

Track the share of requests using each mode and reason together. Normal template use is not degraded traffic. If templates replace generation for 30 percent of requests for two days, the model incident is still visible even though those requests returned successfully.

Test the degraded paths on purpose

Force each dependency outcome with test doubles or fault injection. Verify that:

  1. Unaffected product capabilities still work.
  2. The application never presents a reduced result as the normal one.
  3. A queued job and any required dispatch records are durable before the endpoint accepts it. Concurrent or repeated submissions with the same key do not create another job.
  4. The worker revalidates ticket, policy, and access assumptions before execution.
  5. A relevant ticket or policy change, or loss of required permissions, during generation prevents synchronous and queued drafts from being presented as current.
  6. Duplicate deliveries and concurrent workers do not publish duplicate local results. For external effects that cannot be deduplicated, test the uncertain-outcome policy.
  7. Status polling, cancellation, and result retrieval enforce current authorization, including after access is revoked.
  8. Stale results are rejected after a relevant input or policy change, revoked access, or loss of required permissions.
  9. Recovery returns the flow to normal mode without serving old results indefinitely.

Also test combinations. The model may be unavailable while the queue is full. Retrieval may recover after a job was accepted but before it runs. A previous result may exist after the user’s access was revoked.

These cases are where a fallback can become a data leak or a false promise. When the model recovers, release accumulated jobs at a bounded rate rather than sending the whole backlog at once.

When to continue and when to stop

Design a degraded mode when the AI capability supports a larger user job and the product can preserve useful, honest behavior without it. Good candidates include drafting assistance, optional summarization, enrichment, recommendations, and long-running analysis that can finish later.

Return an explicit unavailable result when every substitute would violate the feature’s promise, required current evidence cannot be obtained, authorization cannot be confirmed, or delayed output would no longer help. Stopping one capability can be the behavior that keeps the rest of the product trustworthy.

Do not treat a second model as an invisible replacement. Another model can change output quality, structured-output behavior, tool use, latency, and cost. That is a separate policy decision, not a transparent implementation detail.

Pick one user flow and write down what remains possible when each dependency fails. Put the resulting behavior in the application contract, show it in the UI, and test it under forced failure.

Sources