Definition
Drift detection compares an AI system’s inputs, outputs, evaluated quality, or operational metrics with a baseline to find changes worth investigating. For production traffic, that baseline can be measurements from an earlier period. For a fixed set of reviewed evaluation cases, it is the recorded results of a previous run. Production data can reveal a new request mix. Rerunning fixed cases can reveal changes in output or measured quality. Compare similar time windows and request groups to help distinguish traffic shifts from quality changes.
The measures depend on the feature. They might cover the kinds of requests users send, properties of generated answers, reviewed answer quality, or operational behavior such as latency and tool failure rates. Detecting a shift tells you where to look. It does not identify the cause.
Simple example
A support assistant usually receives order-status questions. After a return policy update, refund questions become more common. A weekly comparison shows the input mix changing and a smaller share of answers citing a policy document. For requests governed by the new policy, reviewers find answers citing the superseded version. Retrieval logs show some requests reaching outdated documents.
The team reviews refund questions separately and checks whether retrieval had access to the current policy. Latency and tool failure rates are steady. The outdated documents in the logs make document freshness and retrieval the next things to investigate.
Why it matters
A fixed test set can stay green while production traffic moves beyond the cases it covers. Drift detection helps teams notice that gap and decide which live cases deserve review or a place in the evaluation set. It can also expose a gradual rise in latency or tool failures before those changes appear in a release comparison.
One important nuance
Drift is a change, not necessarily a regression. Seasonal traffic or a new customer group may change inputs while the system still performs well. Conversely, a stable overall quality score can hide worse results for one type of request. Small differences may also be sampling noise, especially in a small request group. When a policy changes, judge each request under the applicable version. Record which scoring rules each run used. If the rules differ, the scores may not be directly comparable. Inspect examples before changing the system.