Table of Contents
An IChatClient decorator applies a shared rule to every call through a client. It can reject an oversized request or record a model call. The feature still chooses the prompt and evidence, then checks the answer.
That division is easy to lose. A support service starts with one logging statement. Later it gets a retry loop and a cache lookup. Soon every feature has its own model pipeline, or one large wrapper owns decisions it has no context to make.
The example assumes you already use IChatClient and know the decorator pattern.
Put shared mechanics around the client
Consider a support answer service. It builds messages from a question and approved source material, then asks the model for a draft. The service should own the prompt, source selection, authorization, and response validation. Those decisions vary by feature and caller.
The chat-client pipeline can own behavior that is the same for all calls using that registration. Telemetry belongs there. A fixed request guard can belong there too. The provider adapter remains at the inside of the pipeline:
SupportAnswerService
-> IChatClient
-> request guard
-> telemetry
-> provider client
The service receives one IChatClient, with no direct calls to the guard or telemetry code. Features that need different limits or data-handling rules need separate registrations. Do not infer feature identity from prompt text.
Write a decorator that covers both call shapes
DelegatingChatClient forwards calls to an inner IChatClient. Override the operations where the behavior belongs. Here, a guard rejects a conversation with more than 32 messages at this point in the pipeline. This is a message-count limit, not a token budget. A single large message can still exceed the model’s context window, and an inner decorator could add more messages after this check.
using Microsoft.Extensions.AI;
public sealed class MessageCountLimitChatClient(
IChatClient innerClient,
int maximumMessages) : DelegatingChatClient(innerClient)
{
private readonly int _maximumMessages = maximumMessages is > 0 and < int.MaxValue
? maximumMessages
: throw new ArgumentOutOfRangeException(nameof(maximumMessages));
public override Task<ChatResponse> GetResponseAsync(
IEnumerable<ChatMessage> messages,
ChatOptions? options = null,
CancellationToken cancellationToken = default)
{
ChatMessage[] checkedMessages = Check(messages);
return base.GetResponseAsync(checkedMessages, options, cancellationToken);
}
public override IAsyncEnumerable<ChatResponseUpdate> GetStreamingResponseAsync(
IEnumerable<ChatMessage> messages,
ChatOptions? options = null,
CancellationToken cancellationToken = default)
{
ChatMessage[] checkedMessages = Check(messages);
return base.GetStreamingResponseAsync(
checkedMessages, options, cancellationToken);
}
private ChatMessage[] Check(IEnumerable<ChatMessage> messages)
{
ArgumentNullException.ThrowIfNull(messages);
ChatMessage[] snapshot = messages.Take(_maximumMessages + 1).ToArray();
if (snapshot.Length > _maximumMessages)
{
throw new ArgumentException(
$"A chat request may contain at most {_maximumMessages} messages.",
nameof(messages));
}
return snapshot;
}
}
The guard covers complete and streaming calls. Overriding only GetResponseAsync would leave streaming unguarded. The streaming method checks the messages, then returns the inner stream with the cancellation token. Streaming still uses memory: the OpenTelemetryChatClient registered below retains updates to assemble a final response for telemetry.
Check reads at most one message beyond the limit and materializes the accepted sequence once. The constructor excludes int.MaxValue so _maximumMessages + 1 cannot overflow. The guard does not clone individual messages. Callers should still create messages and ChatOptions for each operation because an IChatClient may mutate them.
The guard does not trim history or choose source material. Those choices can change the answer and belong with the feature.
Compose it once at startup
The provider-specific client is created at the composition root. The example uses Ollama to keep the registration short. The same wrapper can surround another IChatClient implementation. I compiled these snippets with .NET 10, Microsoft.Extensions.AI 10.9.0, and OllamaSharp 5.4.30.
AddChatClient(providerClient) returns a ChatClientBuilder. The chained Use calls add decorators to that pipeline.
using Microsoft.Extensions.AI;
using OllamaSharp;
IChatClient providerClient = new OllamaApiClient(
new Uri(builder.Configuration["AI:Endpoint"]!),
builder.Configuration["AI:Model"]!);
builder.Services
.AddChatClient(providerClient)
.Use(inner => new MessageCountLimitChatClient(inner, maximumMessages: 32))
.UseOpenTelemetry(
sourceName: builder.Environment.ApplicationName,
configure: telemetry => telemetry.EnableSensitiveData = false);
Validate the endpoint and model at startup. If the limit comes from configuration, validate it too.
ChatClientBuilder applies factories in reverse when it builds the pipeline. The first registered decorator is outermost, so the guard runs before OpenTelemetry here. A rejected 33-message request never enters OpenTelemetryChatClient. Its chat-client spans and metrics describe calls that passed the guard, not every attempted application operation. If guard rejections matter operationally, count or log them separately without recording the messages.
UseOpenTelemetry adds chat-client tracing and metrics. Register its activity source and meter with your OpenTelemetry provider. Explicit EnableSensitiveData = false prevents this component from capturing raw messages, tool arguments, and results even if OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT is set. Review metadata separately: ChatOptions.ConversationId can still appear on a span, and the identifier itself may be sensitive. Other components need their own privacy settings. The telemetry tip covers host setup.
If DI owns the outer client, avoid separately registering and disposing the inner provider client. See Dependency Injection for AI Components for lifetime details.
Decide whether logging, caching, or retries belong here
Microsoft.Extensions.AI already has UseLogging, UseOpenTelemetry, and UseDistributedCache. Start with those where their behavior fits. A custom decorator is justified when you need a specific rule that the available components do not express.
Logging and telemetry should describe the work without automatically recording message bodies. A prompt may contain customer data, retrieved documents, or tool results. LoggingChatClient logs messages and options when its logger has Trace enabled (those values may be sensitive). Keep that level off in production for this component. Correlation IDs, duration, and usage data when the provider supplies it are usually more useful for operations than a copy of the conversation.
The built-in DistributedCachingChatClient derives its default key from serialized messages, ChatOptions, and configured additional values. It cannot infer a tenant ID, provider deployment, or prompt version absent from those inputs. UseDistributedCache does not make a shared cache safe for tenant data. If caller context affects an answer, decide how to isolate entries before enabling caching. The configured additional values belong to the client, so they cannot carry per-call tenant context on a shared singleton. The underlying CachingChatClient skips caching by default when ChatOptions.ConversationId is set, since that ID may refer to mutable conversation state missing from the messages.
Even a well-scoped key cannot notice every change in permissions, policy documents, or external state. Check authorization in the application service on every operation, whether the answer comes from cache or provider. Plan expiration or invalidation for changing source material. The built-in distributed cache sets no expiration when writing complete or streaming responses. Unless the backing cache applies its own expiration or eviction policy, entries may persist indefinitely.
The built-in JSON serialization does not guarantee a full round trip for every response property. For example, it ignores ChatMessage.RawRepresentation, and values in ChatMessage.AdditionalProperties may deserialize as JsonElement. Check what your feature needs after a cache hit and how long sensitive output may remain in storage. A hit can also skip inner telemetry and provider calls, depending on decorator order. The HybridCache tip covers stampedes, which are a separate concern.
Retries are even less automatic. A failed request may already have consumed tokens. A tool-call loop may have performed a side effect before the failure reached the decorator. Streaming may have emitted text that a caller already displayed. Blindly replaying the outer call can duplicate work or produce a second, inconsistent answer. Retry only a known transient failure, within a bounded time and attempt budget, and only where replay is safe. The retry article explains that decision in more detail.
Do not treat these components as a list to install everywhere. The pipeline should be short enough that you can say which layer sees a cache hit, which layer sees a provider attempt, and where an exception goes.
When to use this pattern
Use a decorator for a rule that applies consistently to every call through a particular IChatClient registration. Keep it focused, support the call shapes your application uses, and make its position in the pipeline explicit.
Do not put authorization, prompt selection, evidence checks, or product fallback behavior into a shared chat wrapper. Those decisions need the caller and use-case context that a generic decorator should not have. If the feature needs to explain why an answer was rejected or unavailable, return that outcome from the application service.
Related reading
- Provider Independence from Day One
- Designing an AI Service Layer
- Use OpenTelemetry Before You Need Production Debugging
Sources
- Microsoft Learn:
DelegatingChatClient - Microsoft Learn:
AddChatClientregistration - Microsoft Learn:
IChatClientconcurrency and argument contract - Microsoft Learn:
OpenTelemetryChatClient.EnableSensitiveData - Microsoft Learn:
LoggingChatClientandTracedata - .NET source:
ChatClientBuilderconstruction order - .NET source:
OpenTelemetryChatClientstreaming updates - .NET source:
CachingChatClientconversation-state guard - .NET source:
DistributedCachingChatClientkey, serialization, and write path