A fallback model needs to pass the feature’s acceptance checks before it receives production traffic. Sharing an API with the primary model makes integration easier. You still need evidence that it can do the job.

Callers need to know whether they received an accepted routing decision or a suggestion that requires review. If the fallback only supports suggestion mode, the application must return that reduced mode explicitly. Valid JSON does not tell the caller which promise the feature kept.

I would evaluate the fallback as a separate candidate, then decide when the application may use it. Finding out whether it works during an outage leaves little room to fix a bad decision.

A successful response can still break the contract

Consider a ticket router that selects Billing, TechnicalSupport, or AccountAccess. It must request human review when the ticket is ambiguous. The model proposes a decision. Application code checks the response structure and domain constraints, then applies routing policy before any queue assignment.

Consider this ticket:

I cannot access my account to download the invoice. Please help me sign in.

A candidate fallback could return:

{
  "queue": "Billing",
  "requiresHumanReview": false,
  "reason": "The customer needs an invoice."
}

The response passes the structural checks: valid JSON, an allowed queue, required fields, and a reason within the length limit. It still misses the immediate account-access problem. For this example, a reviewer would label the expected queue as AccountAccess.

This is a hypothetical case. No particular model produced it. It shows how a fallback can pass structural and domain validation while assigning tickets incorrectly or suppressing review on ambiguous cases.

The same problem appears in reply drafting. A model can preserve the response shape while omitting a policy exception or making an unsupported commitment. Those are feature failures even if the HTTP request succeeds.

Qualify the complete fallback path

In .NET, IChatClient provides a shared abstraction for chat requests and streaming. Microsoft.Extensions.AI.Abstractions also defines common content types for function calls and results. A pipeline component such as FunctionInvokingChatClient adds automatic tool invocation. Those abstractions help contain SDK dependencies. They do not make every backend satisfy the same feature contract.

Qualify the combination you will actually deploy. Record the model identifier and pin a version or snapshot where the backend allows it. Include the deployment, provider integration, prompt, response schema, and relevant settings in that record. If the fallback needs different instructions, version and evaluate those instructions with it.

A mutable model alias can keep the same configured name while the served model changes. Record observed model metadata when available, and decide how an upstream change triggers requalification. Your qualification evidence applies to the configuration you tested. It can become stale after a change underneath the alias.

For the ticket router, the qualification record should answer these questions:

ConcernEvidence needed before enabling fallback
Inputs and contextRealistic ticket lengths and supported languages fit without dropping required information
Structured outputThe exact schema works through the configured integration. Missing fields, refusals, and incomplete responses have defined outcomes
Routing qualityReviewed cases cover queue selection and ambiguity, including tickets that mention several problems
Tool behaviorIf tools are part of the feature, the candidate selects permitted tools with correct arguments and handles tool failures
Latency and costThe complete candidate path fits its allowed time and spend, including validation and any permitted repair
Data boundaryThe destination is approved for the request’s data, tenant, region, and retention requirements

Reject some candidates before measuring quality. If a destination cannot receive the ticket data, its routing score is irrelevant. The same applies if it cannot accept the required input.

Define acceptable error rates before evaluation, along with requirements for critical cases. Specify which cases must go to review and which failures disqualify a candidate. A good overall score can hide errors on account-access tickets that have a different consequence from a misrouted billing question.

Report the number of cases behind each rate and how you selected them. Break results down by relevant groups, such as language, ambiguous tickets, and account-access requests. Consider sampling uncertainty before treating a small passing set as evidence for production behavior. A candidate should not pass merely because the evaluation contains mostly easy billing tickets.

Check what the fallback path can actually bypass. It might select another model or deployment behind the same endpoint, another region, or a different provider. Map the dependencies it shares with the primary path, including credentials, network access, and capacity limits. Switching models through the same failed gateway will not restore service.

Keep structured-output claims precise

Schema-constrained structured output and asking a model to write JSON are different mechanisms. OpenAI’s structured-output documentation distinguishes JSON mode from schema adherence, documents a supported subset of JSON Schema, and describes refusal handling. Check the corresponding documentation and integration behavior for each candidate.

If the primary path relies on schema-constrained structured output and the fallback uses prompt-guided JSON, that difference belongs in the qualification record. The fallback might still meet the feature’s acceptance criteria with application validation. A single successful smoke test cannot establish that.

Validate required fields and domain rules on every path. Treat an incomplete response or refusal as a defined outcome rather than forcing it into the success DTO. Schema compliance establishes properties of the response shape. Evaluation provides evidence about routing quality, and trusted system controls decide whether the application may act on a proposal. Ordinary validation code does not prove that the model chose the correct queue.

Define which failures permit a switch

A catch-all handler can send a cancelled request to another model or conceal a bug in the application. Neither case calls for another generation attempt.

For this router, I would start with a narrow policy:

Primary outcomeFallback policy
Temporary model service unavailabilityPermit one qualified alternate if the request still has enough budget
Provider throttlingPermit the alternate only if it has available capacity and the request remains within application admission limits
Attempt timeoutPermit the alternate only if the overall deadline remains useful and the operation is safe to repeat
Caller cancellation or overall deadline exhaustedStop
Invalid credentials, unsupported options, or an application exceptionSurface the integration failure. Do not route it through generic fallback
Safety refusal or a denied application actionPreserve the safety or authorization outcome. Do not search for a model that will comply
Invalid model outputReturn the defined invalid-result or review outcome. Any repair or alternate-generation policy needs separate qualification and budget

Classify failures in the provider adapter and expose stable categories to the application. Avoid matching arbitrary exception text. The table above is a proposed feature policy, not a universal mapping from HTTP status to fallback.

Automatic fallback should also preserve required evidence. If the ticket router cannot obtain context required by its contract, the application must not treat generated content as a substitute for that missing context. Graceful Degradation for AI Features covers the product alternatives when generation cannot safely continue.

Keep fallback inside the request budget

Fallback spends another model attempt and more input tokens from the original execution budget. The primary attempt may already have incurred cost, even when the application received no usable response. Keep that spend accounted for. Missing usage data after a timeout is not evidence that the attempt was free.

Take a request with an eight-second deadline. If the primary attempt consumes six seconds, the fallback has at most two seconds left, including validation and result handling. Giving it a fresh eight-second timeout would let the operation outlive the caller’s deadline.

Before each attempt, verify that the candidate fits within the remaining time and spend budget, including completion work. Estimate spend from the input, configured output limit, and applicable pricing, including any other billable work on that path. Actual completion time and billed usage may remain unknown until later. If the candidate does not fit the admission policy, return the chosen unavailable or review outcome. Reuse the caller’s cancellation signal and the overall deadline. An attempt timeout may be shorter.

For this example, allow at most two outbound model attempts in total. The intended path is one primary attempt followed by one fallback attempt if needed. Configure, disable, or account for SDK retries explicitly so they cannot add attempts outside the shared deadline and attempt budget. Under this policy, if the provider client performs a second outbound model attempt, it consumes the remaining slot. The application cannot then start the fallback. A separately qualified repair call also consumes that slot instead of creating a third. AI Systems Need Runtime Budgets develops that accounting in more detail.

When an SDK retry layer cannot participate in that accounting, disable it for this policy. Another design can give client invocations and underlying transport attempts separate limits, provided it counts both and keeps them inside the same execution budget. The policy must make those units clear. Counting only calls to IChatClient can miss repeated requests inside a provider client.

The fallback also needs a capacity limit. It must not inherit an unlimited stream of requests during a primary outage. If it has no capacity, stop or use another approved degradation mode.

Stop at uncertain side effects and partial output

The ticket-routing model call proposes a decision without assigning the ticket, so the application can repeat generation before committing anything. Tool-enabled workflows need a stricter boundary.

If the primary path invoked a tool that saved a draft, a timeout does not prove that the save failed. Establish the outcome or use an idempotent write contract before repeating work through another model. Resume from known state and avoid repeating completed actions. Keep authorization and approval enforcement outside the model, in trusted system controls, on both paths. Depending on the architecture, those controls may live in the application, a tool service, a policy engine, or the downstream system.

Streaming creates another boundary. Once the application has shown part of an answer, silently appending a fallback model’s continuation can mix incompatible answers. For this policy, automatic switching is allowed only before visible output and before any unresolved side effect. After that boundary, return an interrupted outcome or offer an explicit restart with known state.

What the .NET failover APIs handle

Microsoft.Extensions.AI 10.9.0 introduced experimental routing and failover APIs, marked with diagnostic MEAI001. RoutingChatClient selects a client per request. FailoverChatClient adds the failover loop, and OrderedFailoverChatClient tries clients in a configured order.

The failover loop can select another client when an invocation fails before any streaming update reaches its caller. After an update has been exposed, the failure propagates without a switch. Request cancellation also stops reselection. MaximumAttemptsPerRequest limits client invocations. It does not count retries hidden inside a provider SDK. That caller boundary can be earlier than visible UI text: a streaming update may reach a pipeline component before the page displays it.

FailoverChatClient can try another client after an invocation exception without first deciding whether the failure is transient. If the request has not been cancelled, no streaming output has been committed, and the attempt limit allows it, another selection may follow. Invalid credentials or unsupported options can therefore trigger a switch too. OrderedFailoverChatClient supplies the order. It does not implement the failure categories in the earlier table.

Block the next client invocation when the failure category is ineligible. A custom FailoverChatClient can inspect attempt updates and keep policy state for the request to enforce that rule. Alternatively, a provider-aware application policy can own dispatch and decide whether to invoke the alternate at all. Check the remaining budget and tool state before dispatch too. The application still needs evidence that the candidate can do the job and that its destination may receive the data. Check the experimental API behavior against the package version you deploy.

Pipeline placement changes tool behavior: routing around FunctionInvokingChatClient selects a client for the whole tool-loop attempt, while routing inside it selects for individual model turns. Qualify the pipeline arrangement and failover boundary you actually deploy.

Make the actual mode visible

A qualified alternate may meet the normal contract. Another candidate may only be suitable for suggestions that require review. The application result needs to tell callers which behavior they received.

For the router, a reduced fallback mode could display:

Automatic routing is temporarily unavailable. Review the suggested queue before assigning the ticket.

In that mode, application code requires review regardless of the model’s requiresHumanReview value. The caller must not infer the mode from that model-generated field.

Keep the model’s proposal separate from the result the application returns:

public enum TicketQueue
{
    Unknown = 0,
    Billing = 1,
    TechnicalSupport = 2,
    AccountAccess = 3
}

public abstract record RoutingResult
{
    public sealed record Accepted(TicketQueue Queue) : RoutingResult;
    public sealed record ReviewRequired(TicketQueue SuggestedQueue)
        : RoutingResult;
    public sealed record Unavailable : RoutingResult;
}

public static class TicketRoutingPolicy
{
    public static RoutingResult Apply(
        TicketQueue proposedQueue,
        bool modelRequestedReview,
        bool backendRequiresReview) =>
        backendRequiresReview || modelRequestedReview
            ? new RoutingResult.ReviewRequired(proposedQueue)
            : new RoutingResult.Accepted(proposedQueue);
}

Before calling Apply, check the response structure and domain constraints, including the allowed queue and the required boolean review field. The name proposedQueue is deliberate: checking those constraints does not establish that the queue is correct. The application’s qualified backend policy supplies backendRequiresReview. When that flag is true, the result is always ReviewRequired, even if the model requested automatic assignment. For a fully qualified primary or fallback backend, Accepted is available when policy permits automatic routing and the model has not requested review. Return new RoutingResult.Unavailable() when no usable proposal exists.

Keep missing fields visible at the model boundary. With TicketQueue? Queue and bool? RequiresHumanReview in the response DTO, an omitted field remains null, while an explicit review value of false remains false. Alternatively, enforce required-field presence during deserialization. Reject missing fields, Unknown, queues outside the three allowed values, and reasons that exceed the length limit before calling Apply. default(TicketQueue) is Unknown, which validation must reject. Apply keeps its non-nullable parameters after the parsing boundary has checked field presence and domain constraints.

The assignment path must accept automatic routing only for Accepted. A ReviewRequired result displays the suggested queue and waits for a separate, authorized human decision. These result types carry the policy across the service boundary. The model cannot grant permission to assign a ticket by returning requiresHumanReview: false.

Accepted means the application accepted the proposal for automatic routing under its policy. It does not certify semantic correctness or replace authorization at the assignment boundary. The example represents the outcome decision, not the validation or permission checks themselves.

If a fallback keeps the user promise unchanged, the feature contract alone does not require a provider banner on every response. Check contractual, regulatory, and customer provenance requirements separately. Those may still require disclosure. Record the actual backend internally, and expose provenance where required by those obligations or the API contract. Disclose changes that affect what the user can rely on, such as mandatory review, a smaller accepted input, or unavailable actions.

Record the feature mode and switch reason with the attempt count, elapsed time, and accounted usage. Include the qualification version and configured model identity, plus the observed identity when the provider returns it. Avoid putting ticket text or generated reasons into routine telemetry. Report fallback traffic separately so successful alternate responses do not hide a persistent primary incident.

Carry the mode and provenance into stored or cached results too. A review-only suggestion must not later appear as an accepted routing decision because a cache discarded the metadata.

Evaluate the fallback before the outage

Run the primary and fallback paths against the same frozen, reviewed evaluation set. Each path uses its own recorded configuration. Compare incorrect queue assignments and missed review cases separately from parse failures. Include language and input-length cases the feature actually accepts.

Keep development cases separate from the qualification set. Use development cases to tune prompts and settings, then freeze each candidate before the qualification run. If you repeatedly adjust a candidate in response to the qualification cases, those cases become part of development. Use a fresh held-out set for the next qualification decision rather than treating a tuned passing score as independent evidence.

Use the acceptance thresholds you set before inspecting the results. Report how many cases failed out of how many evaluated, broken down by relevant groups and failure severity. Repeating a generation on the same inputs helps reveal model variability. You still need varied cases, and neither check removes uncertainty about production traffic. Passing the critical cases in this set cannot prove that the model will avoid the same mistake on every future ticket.

For a reduced mode, evaluate that mode’s promise as well. A suggestion may be useful with mandatory review even when it is unsuitable for automatic assignment. Review does not make every suggestion acceptable. Define what would make it misleading enough to reject entirely.

Then force the primary path to fail. Verify the switch reason, remaining budget, returned mode, and downstream behavior. Include cancellation before fallback, an exhausted deadline, invalid fallback output, and an unavailable alternate. For tool or streaming features, test the boundary after visible output or a possible write.

A passing smoke test establishes that a path can work. The broader evaluation, with varied cases and repeat runs where needed, provides evidence about how often it satisfies the task. Requalify when the model version, prompt, schema, integration, or relevant settings change. Keep the evidence with that configuration. What Should You Actually Evaluate? explains how to choose those acceptance checks from the feature contract.

When to use fallback models

Use a fallback when a qualified alternate can preserve the feature’s contract, the failed dependency can be bypassed, and another attempt fits the remaining budget. Bounded classification or extraction can be good candidates because the application can check structural and domain constraints and gate downstream actions before committing them. Evaluation must still establish whether the candidate’s task quality is acceptable.

A narrower fallback can also help when the product names the changed behavior and enforces it. A reviewed routing suggestion is useful only if the assignment path actually requires review.

When not to use fallback models

Stop when the alternate cannot satisfy required capabilities or data boundaries, the request is cancelled, or insufficient execution time remains. Do not switch to bypass a safety decision or repeat an uncertain write.

Before enabling fallback for one feature, run its reviewed cases through the alternate path. Then force a primary failure and inspect the result at the caller, including whether the application enforces review. Keep the fallback disabled until those checks pass. If no acceptable alternate exists, choose an explicit unavailable or non-AI path.

Sources