The same test questions scored 84% at first and 51% later: model drift in AI
Same questions, worse answers

TL;DR

  • What is model drift in AI? It is when an AI system that worked well at launch slowly gets worse, because something changed after it was tested.
  • The questions, the right answers or something behind the scenes can change, such as a provider updating its model.
  • You catch it by re-running the same fixed test regularly, once you know how much scores wobble on their own.

Key Takeaways

  • Drift happens after launch, not during training. A model can pass every test on day one and still get worse in use.
  • With AI services, the model itself can change under you. In one study, a popular model’s accuracy on the same maths questions fell from 84% to 51% in three months.
  • Know the normal wobble before calling anything drift. On our 50-question tests, gaps smaller than 0.27 to 0.45 points out of 10 can’t be told apart from noise.
  • The measuring tool can drift too. We re-check our AI grader against 8 answers with known scores every 25 answers, and stop if any score moves.

A model that passed every test at launch can be worse six months later without anyone touching it. What is model drift in AI, and why does it happen to systems nobody changed? Usually because something around the model did change: the data coming in, what counts as a correct answer, or the model behind an API.

We run Groundtruth, Eigenform’s open source dynamic benchmark, and re-grade models on fixed question sets. That puts us on the measurement side: telling a real change in a model from a change in the test, the grader or noise.

What is model drift in AI?

Model drift is the decline in a deployed AI model’s performance over time, as the data it sees, the relationships it learned or the system around it change after training. The model’s weights may not change at all; its world does.

Some sources reserve the term for concept drift. Here, however, it is the umbrella for any loss of performance after deployment, with data drift as one of its causes.

What is drift in AI systems built on large language models? Here, retrieval or memory that surfaces different examples changes the system’s effective knowledge without a single weight moving.

Model drift vs data drift: the types of drift

Model drift vs data drift is the most common confusion in AI monitoring. Data drift is a change in the inputs; model drift is the drop in performance that may or may not follow it.

TypeWhat changesExample in exploration workHow it shows up
Data driftThe distribution of inputsA gold-report assistant is asked about copper elsewhereUnfamiliar names, units and formats
Concept driftThe link between input and correct answerA new discovery changes what is known about an areaSame question, outdated answer
Label driftHow often each outcome occursWork moves to greenfield groundThe share of positive calls shifts
Upstream driftA component the model depends onA lab’s new detection limit; a provider’s model updateA step change after the switch

Data drift (covariate shift)

The inputs change while the correct mapping stays the same, so the model faces cases it rarely saw in training. With the same eight models and the same judge, moving to a different document set reordered the lower half of the ranking. One model was the lowest of its middle-tier group on the Technical set (7.34) and the highest of that group on WAMEX (8.46).

Concept drift

The question stays the same, but the correct answer changes. In exploration, a new discovery or a revised interpretation changes what is true about an area, while a model trained earlier still gives the old answer.

For a language model, every fact that changes after its training cutoff is concept drift in miniature, because the weights stay where pre-training left them.

Historical sources raise the same problem from the other side. In our benchmark, each 1920s report is the authority for what it claims, and no question asks a model to correct it against present-day understanding.

Label drift

How often each outcome occurs changes, while what each outcome looks like stays the same. When work moves from ground near a known deposit to greenfield ground, far fewer samples are anomalous, and a model calibrated on the old mix raises more false alarms. Lipton and colleagues formalise this as label shift.

Upstream and feature drift

Something the model depends on changes: a data source, a parser, a lab method, or the model behind an API. Nothing in your own code changed, which makes it hard to see.

For example, re-extracting a corpus moves every line boundary, so every evidence anchor has to be verified again.

Assay data drifts the same way. A lab that lowers its detection limit, or starts reporting <0.005 instead of a number, shifts the column although the rock has not changed.

Why does a model drift?

Every cause of model drift in AI is a change after the model was tested:

CauseUsual drift type
Exploration moves to a new district or commodityData drift
A discovery or reinterpretation after the training cutoffConcept drift
Work moves from near-mine to greenfield groundLabel drift
A lab changes its method, units or detection limitFeature drift
A provider updates or retires a modelUpstream drift
The system learns from feedback or its own outputAny of them

Above all, provider updates deserve attention. In one widely cited study, GPT-4’s accuracy at identifying prime numbers fell from 84% in March 2023 to 51% in June 2023, on the same questions. Pinning a dated model version helps; however, pinned versions are eventually retired.

Moreover, systems that keep learning add drift of their own. Self-training without an outside check degrades through reward hacking and semantic drift: each round is judged against the previous one, so a model drifting toward a narrower distribution goes unnoticed.

How to detect model drift: data drift detection and performance checks

Detecting model drift in an AI system means comparing current behaviour with a reference. Checks on the inputs need no labels; checks on the outputs do.

Statistical distribution tests

Data drift detection starts with the inputs. The population stability index (PSI) and Kullback–Leibler (KL) divergence compare a feature’s distribution today with its distribution at training time: this quarter’s gold assays against the training set, for example. The two-sample Kolmogorov–Smirnov (KS) test asks whether two samples come from the same distribution.

However, each has limits. PSI thresholds such as 0.1 and 0.25 are rules of thumb, not calibrated error rates, and the KS test checks one variable at a time and flags trivial differences on large samples. For LLM inputs, the features are usually embeddings, prompt lengths or topics, and tools such as Evidently choose a test by column type and sample size.

Performance monitoring against fresh ground truth

Input tests say that something changed; however, only fresh labels say whether it matters. That needs questions with verified answers, the ground truth data for AI that outputs are scored against, refreshed to reflect today’s cases.

For LLM systems, the practical version is a frozen benchmark. When a vendor proposes a model that is 40% cheaper, re-run the same frozen benchmark on your own tasks before switching.

Re-running an unchanged reference system also tells you whether a score moved because a model changed or because the benchmark itself quietly drifted.

Data drift monitoring: dashboards and alert thresholds

A threshold set inside the noise fires on nothing, so measure the noise floor first. On the 50-question sets behind the Groundtruth leaderboard, the smallest gap we could distinguish from noise was 0.27 to 0.45 points out of 10.

In addition, the grader adds noise of its own. Our judge is not perfectly deterministic even at temperature 0, so a re-run can move a close comparison.

Likewise, monitor the grader. On our WAMEX benchmark, a standardisation set of eight answers with known scores is re-scored at the start of every grading run and after every 25 responses. If a pass/fail decision changes or a fixed award moves, grading stops until the rubric is recalibrated.

Human review sampling as a backstop

Still, automated checks miss failures that look valid. In one of our runs, the harness took the question for the LLM-as-a-judge from the answer file, so a mismatched field made it grade against the wrong question. Its output looked normal and the log showed only a warning; only reading real transcripts can catch that.

Finally, route uncertain cases to people. Our judge flags answers it cannot verify, and flagged questions are listed after every run for a person to check.

Responding to drift

Three responses close the loop:

  • Re-evaluate on events, not only on a calendar: a model release, a vendor switch, a change to retrieval or prompts, or new data.
  • Keep a rollback path. Without snapshots, versioned prompts and weight checkpoints, improvement cannot be separated from drift.
  • Retrain on recent data. In our self-training experiments, retraining on a sliding window of recent data beat both retraining on everything and retraining on only the best runs.

FAQs

Can a model provider’s update cause model drift?

Yes, and it is a common source of model drift in AI products. A hosted model behind an API alias can change with nothing changing on your side, so record which version produced each result and re-run a frozen benchmark after every change.

What is the difference between model drift and data drift?

Data drift is a change in the inputs a model receives. Model drift is the resulting loss of performance. Data drift can occur without model drift if the model handles the new inputs well, and model drift can occur without data drift when the correct answers change.

What causes a model to drift?

Changes after deployment: new regions or commodities, discoveries after the training cutoff, a different mix of outcomes, and upstream changes such as a lab’s new detection limit or a provider’s model update. Systems that learn from feedback or their own output can drift too.

How do you detect model drift?

Compare current behaviour with a fixed reference. Tests such as PSI or Kolmogorov–Smirnov flag changes in the inputs; scoring outputs against fresh ground truth, or re-running a frozen benchmark, shows whether performance fell. Set alert thresholds above the noise between repeated runs.

How often should models be monitored for drift?

Monitor inputs continuously, and re-test outputs after every event that could change behaviour: a model or provider update, a retrieval or prompt change, or new data. In long graded evaluations, re-check the grader too; we re-score a fixed set of known answers every 25 responses.