Model drift is the moment a validated AI model quietly stops being the model you validated. The code has not changed and nobody has pushed a release, yet the predictions have moved because the world underneath them moved. Most guides treat model drift as an engineering nuisance to be fixed with a retraining job. Under the EU AI Act, the GDPR accuracy principle and banking supervision, it is something more serious: the point where a system that passed its conformity assessment can stop meeting the accuracy it declared, without anyone signing off on the change.

Key takeaways
- Model drift is the decline in a model’s predictive performance after deployment, caused by changes in input data (data drift) or in the relationship between inputs and outcomes (concept drift).
- Article 15(1) of the EU AI Act requires high-risk systems to perform consistently throughout their lifecycle, and Article 15(3) makes providers declare accuracy levels. Undetected drift turns that declaration into a false statement.
- Article 72 makes post-market monitoring mandatory for high-risk providers. The Digital Omnibus replaced the monitoring-plan implementing act with Commission guidance and a template due by 2 September 2027.
- Retraining inside a change envelope documented at the initial conformity assessment is not a substantial modification (Article 43(4)). Retraining outside it can be.
- For LLMs, drift often comes from the vendor. The same API name can serve different behaviour months apart, so pinning versions and contracting for change notice are drift controls.
What is model drift?
Model drift, also called model decay, is the gradual or sudden loss of predictive quality in a machine learning model once it runs on live data. IBM and Stanford HAI define it the same way: the model was fitted to one snapshot of reality, and reality kept moving. Three mechanisms sit behind the term, and they call for different responses.
- Data drift (covariate shift). The statistical distribution of the inputs changes. A fraud model trained on card-present transactions starts seeing mostly wallet payments. The rules linking inputs to outcomes may still hold, but the model is now operating in regions of the input space it barely saw during training.
- Concept drift. The relationship between inputs and the target changes. The same income and debt profile that signalled low credit risk in a stable economy signals higher risk during an inflation shock. The inputs look familiar; their meaning has shifted.
- Label drift (prior probability shift). The base rate of the outcome changes. If fraud prevalence doubles, a model calibrated on the old rate will under-alert even if nothing else moved.
Model drift can also be sudden (a pandemic, a regulatory change, a new product launch), gradual (slow demographic change), seasonal (retail demand) or recurring. Upstream pipeline changes, such as a renamed field or a unit switched from euros to cents, produce symptoms identical to drift but are really data-quality defects. Telling them apart is the first step of any response.
Model drift vs data drift vs concept drift
<table header-row=”true”> <tr> <td>Term</td> <td>What changes</td> <td>Typical signal</td> <td>First response</td> </tr> <tr> <td>Data drift</td> <td>Input distribution</td> <td>PSI or KS test on features</td> <td>Check the pipeline, then assess impact</td> </tr> <tr> <td>Concept drift</td> <td>Input-to-outcome relationship</td> <td>Accuracy drop on labelled outcomes</td> <td>Re-validate, likely retrain</td> </tr> <tr> <td>Label drift</td> <td>Outcome base rate</td> <td>Shift in predicted vs actual rates</td> <td>Recalibrate thresholds</td> </tr> <tr> <td>Model drift</td> <td>Overall performance (the effect)</td> <td>Any of the above</td> <td>Triage the cause before acting</td> </tr> </table> In short, data drift and concept drift are causes; model drift is the observable effect on performance.
Why model drift is a compliance problem, not only an MLOps one
The top-ranking explanations of model drift stop at dashboards and retraining schedules. That misses what a regulator or auditor will ask, which is not “did you retrain?” but “how did you know, what did you decide, and where is the record?” Three legal anchors turn model drift into a compliance event. The accuracy you declared. Article 15 of the EU AI Act requires high-risk AI systems to “perform consistently in those respects throughout their lifecycle”, and paragraph 3 requires the levels of accuracy and the relevant metrics to be declared in the instructions for use. A provider that declared 94% recall at conformity assessment and now runs at 86% because of drift is shipping a product whose documentation no longer describes it. The same article asks that systems which keep learning after deployment address feedback loops, where biased outputs become tomorrow’s training data. The accuracy principle in data protection. Where a model makes or supports decisions about people, degraded predictions are inaccurate personal data in the sense of Article 5(1)(d) GDPR. The European Data Protection Supervisor made this explicit in its November 2025 guidance on AI risk management, which lists “inaccurate output due to data drift” as a named risk. Its example is a credit scoring model trained during a stable economy and then used during a crisis with sharp changes in inflation and unemployment. The recommended measures are drift detection, data quality monitoring, retraining on a fixed schedule or on detected drift, and user feedback channels. Fairness that decays. A model can hold its aggregate accuracy while its error rate for one subgroup climbs, because the subgroup’s data shifted and the majority’s did not. Drift monitoring that tracks only global metrics will miss it. That is why fairness metrics belong in the same monitoring plan, as covered in our guide to AI bias.
What the rules actually require for model drift
No regulation uses the words “model drift” as a defined term, but several require exactly what drift management produces. The table below maps the main regimes. <table header-row=”true”> <tr> <td>Regime</td> <td>Provision</td> <td>What it means for drift</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 15(1), 15(3)</td> <td>Consistent performance over the lifecycle; declared accuracy metrics</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 9(2)(c)</td> <td>Risk management must evaluate risks surfaced by post-market monitoring data</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 12</td> <td>Automatic logging that makes post-deployment monitoring possible</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 26(5)</td> <td>Deployers monitor operation and inform the provider of risks</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 72</td> <td>Providers run a documented post-market monitoring system and plan</td> </tr> <tr> <td>EU AI Act</td> <td>Art. 73</td> <td>Serious incidents reported within 15 days, or 2 or 10 days in specific cases</td> </tr> <tr> <td>NIST AI RMF</td> <td>MEASURE 2.4, MANAGE 4.1</td> <td>Production behaviour monitored against pre-deployment results</td> </tr> <tr> <td>ISO/IEC 42001</td> <td>Clause 9.1, Annex A operation and monitoring control</td> <td>Monitoring, measurement and evaluation of AI system performance</td> </tr> <tr> <td>US banking</td> <td>SR 26-2 (April 2026)</td> <td>Risk-based model risk management, replaces SR 11-7</td> </tr> </table> Post-market monitoring under Article 72. Providers of high-risk systems must “establish and document a post-market monitoring system” that actively collects and analyses performance data over the system’s lifetime (Article 72). The plan sits in the Annex IV technical documentation. The original text required a Commission implementing act with a template by 2 February 2026. Regulation (EU) 2026/1744, the Digital Omnibus on AI, replaced that with Commission guidance, including a template, due by 2 September 2027. Stand-alone Annex III high-risk obligations now apply from 2 December 2027 and Annex I embedded systems from 2 August 2028, so the template is expected a few months before the first systems fall in scope. Teams should not wait for it: the elements of a credible plan are already clear. NIST AI RMF. MEASURE 2.4 states that the functionality and behaviour of the AI system are monitored in production, and the NIST playbook names drift directly as the reason. It asks teams to compare production metrics with pre-deployment test results and to alert when input or prediction distributions diverge. See our NIST AI RMF guide for the full function map. US banks: a scope gap worth knowing. On 17 April 2026 the Federal Reserve issued SR 26-2, which supersedes SR 11-7, the reference text for model monitoring since 2011, and the OCC issued a matching bulletin. The revised guidance is shorter and more risk-based, and a footnote excludes generative and agentic AI models from its scope, pending a request for information. The practical effect is that the models most exposed to vendor-side drift now sit outside the bank supervisors’ model risk guidance, while conventional credit and fraud models remain inside it. Banks still need a monitoring answer for their LLM use cases; they simply have to build it from their own risk framework. Our article on model risk management covers the wider discipline.
Drift or substantial modification? The retraining decision
Retraining is the standard fix for model drift, and it creates a legal question most technical guides skip. Under Article 3(23) a substantial modification is a change after placing on the market that was “not foreseen or planned in the initial conformity assessment” and that affects compliance or the intended purpose. A substantial modification means a new conformity assessment. Article 43(4) gives providers a way through. For systems that continue to learn, changes to the system and its performance that were pre-determined at the initial conformity assessment and described in the technical documentation are not substantial modifications. The consequence is practical: the retraining envelope has to be written down before launch, not improvised after an alert. A defensible envelope states:
- What may change. Model weights through retraining on the same feature set and the same data sources, for example. Not new features, not a new model family.
- The data that retraining may use. Sources, time windows, quality gates and the representativeness checks from Article 10.
- The acceptance criteria. The minimum accuracy, calibration and subgroup fairness a retrained model must reach before release, tied to the metrics declared under Article 15(3).
- The validation method. Hold-out design, comparison against the live model, sign-off roles.
- The boundary. Which changes fall outside the envelope and trigger a substantial-modification review.
This mirrors the logic of the US Food and Drug Administration’s predetermined change control plan guidance, finalised in December 2024 for AI-enabled medical devices: describe the planned modifications, the method to develop and validate them, and the impact assessment, all up front. Medical device teams subject to both regimes can reuse one change-control design for both.
How to detect model drift
Detection has two layers, and a mature programme runs both. Performance monitoring on ground truth. When outcomes arrive (a loan defaults or not, a claim turns out fraudulent or not), compare predictions with reality and track accuracy, recall, calibration and subgroup error rates against the declared baseline. This is the only direct measure of model drift. Its weakness is delay: credit outcomes can take months. Distribution monitoring as an early warning. Because labels lag, teams watch the inputs and outputs for change:
- Population Stability Index (PSI) compares binned distributions between a reference window and a live window. A widely used rule of thumb reads below 0.1 as stable, 0.1 to 0.25 as moderate shift and above 0.25 as significant shift. These are conventions, not law, and each threshold you adopt should be justified in the monitoring plan.
- Kolmogorov-Smirnov test for continuous features, and chi-square for categorical ones.
- Wasserstein distance and Jensen-Shannon divergence where a magnitude of shift is more useful than a p-value.
- Prediction drift: the distribution of scores or classes the model emits, which moves before labels confirm anything.
Two cautions. First, statistical tests on large volumes flag trivial shifts as significant, so pair them with effect-size thresholds. Second, data drift without performance loss is common. An alert should open an investigation, not trigger an automatic retrain. For how benchmarks feed this baseline, see AI benchmarking.
Model drift in LLMs and third-party models
Large language models drift in a way classic MLOps tooling was not built for: the change often happens on the vendor’s side. In a widely cited study, Chen, Zaharia and Zou compared the March and June 2023 versions of GPT-4 and found its accuracy at identifying prime versus composite numbers fell from 84% to 51%, with instruction-following degrading more broadly. Same product name, three months apart. For a deployer, that makes LLM drift a third-party risk as much as a technical one. The controls:
- Pin model versions where the provider offers dated snapshots, and treat moving aliases as unapproved for high-risk use.
- Run a fixed evaluation set (golden prompts with expected behaviour, refusal tests, format checks) on every version change and on a schedule, and keep the results.
- Contract for change notice: advance notice of model updates and deprecations, and access to release notes. Our vendor due diligence guide lists the clauses.
- Monitor inputs too. User behaviour shifts, and retrieval corpora change under retrieval-augmented systems, both of which move outputs without any vendor update.
A model drift monitoring plan an auditor will accept
Whatever the Commission template eventually looks like, a plan that answers the following will hold up under Article 72, ISO/IEC 42001 and NIST alike.
- Scope and ownership. Which model versions, which use cases, and a named owner accountable for drift decisions. Link each model to its entry in your AI inventory and registry.
- Metrics and baselines. The declared accuracy metrics from the instructions for use, subgroup fairness metrics, and the distribution metrics used as early warnings, each with its baseline value and the dataset it came from.
- Data sources. Where production inputs, predictions and outcomes are logged (Article 12 logs are the natural source), and how deployer feedback reaches the provider.
- Cadence. How often each metric is computed and reviewed, and by whom. Higher-risk and faster-moving domains get shorter cycles.
- Thresholds and escalation. Warning and action levels for each metric, and who is notified at each level.
- Response procedure. The steps from alert to decision, including the substantial-modification check and the incident check.
- Records. Every alert, investigation, decision, retraining run and validation result, retained with the technical documentation. An auditor should be able to reconstruct why the model in production today is the model it is.
The records are the part most teams underinvest in. A dashboard that turned red and then green again proves nothing unless someone can show what was decided in between. This is the ground covered by compliance monitoring for AI systems more broadly.
Responding to model drift: a five-step playbook
- Triage the cause. Rule out pipeline defects first: schema changes, broken joins, unit changes, missing values. A surprising share of “drift” alerts are data-quality bugs.
- Assess impact. Is performance on labelled outcomes below the acceptance criteria? For which subgroups? What decisions has the model influenced since the shift began?
- Contain. Depending on severity, raise the human review rate, tighten decision thresholds, fall back to a previous version or suspend automated decisions. Article 14 human oversight measures should already define these switches.
- Correct. Retrain or recalibrate inside the documented envelope, validate against the acceptance criteria and release through the normal approval path. If the fix falls outside the envelope, open a substantial-modification review before release.
- Report and learn. If drift has led to an outcome that meets the Article 3(49) definition of a serious incident, the Article 73 clock applies: 15 days as the general rule, 10 days where a death is involved and 2 days for widespread infringements or serious disruption of critical infrastructure. Our AI incident reporting guide details the procedure. Either way, feed the finding back into the AI risk management process, as Article 9(2)(c) requires.
FAQ
What is model drift in simple terms? Model drift is when a machine learning model gets worse after deployment because the data it sees, or what that data means, has changed since training. The model itself is unchanged; the world around it has moved. It shows up as falling accuracy, miscalibrated scores or growing errors for particular groups, and it is expected in almost every production system, which is why it has to be monitored rather than assumed away. What is the difference between data drift and model drift? Data drift is a change in the distribution of the model’s inputs. Model drift is the resulting loss of performance. Data drift can happen without model drift, if the model generalises well to the new inputs, and model drift can happen without visible data drift, when concept drift changes the relationship between inputs and outcomes. Monitor both: data drift as an early warning, performance as the confirmation. What is LLM model drift? LLM drift is a change in a language model’s behaviour over time. It can come from the provider updating the model behind the same name, from changes in prompts, user behaviour or retrieved documents, or from fine-tuning. Research on GPT-4 found large performance swings between versions released three months apart. Pin versions, run a fixed evaluation set on every change, and contract for advance notice of updates. Does retraining a model count as a substantial modification under the EU AI Act? Not if the retraining was pre-determined. Article 43(4) states that changes planned by the provider at the initial conformity assessment and documented in the technical documentation are not substantial modifications. Retraining that changes features, data sources or intended purpose, or that falls outside the documented acceptance criteria, may be one and would require a new conformity assessment. How often should you check for model drift? There is no single legal frequency. Set the cadence by risk and by how fast the domain changes: daily or continuous distribution checks for high-volume decisioning, weekly or monthly performance reviews where labels arrive, and a formal review at least annually. Document the rationale in the post-market monitoring plan, since that is what an assessor will test. Is model drift a serious incident under Article 73? Drift on its own is not. It becomes reportable when it directly or indirectly leads to an outcome in the Article 3(49) definition: death or serious harm to health, serious and irreversible disruption of critical infrastructure, infringement of fundamental-rights obligations, or serious harm to property or the environment. A drifting credit or hiring model that produces discriminatory outcomes at scale can meet the fundamental-rights limb.
Conclusion
Model drift is certain; what varies is whether you can prove you handled it. The technical side is well served by existing tools: distribution tests, performance tracking, retraining pipelines. The governance side is where most organisations are exposed, because the EU AI Act, the GDPR and banking supervision all ask the same question in different words: did the model keep doing what you said it would, and can you show it? Answer that with three artefacts written before launch: declared metrics with baselines, a post-market monitoring plan with thresholds and owners, and a retraining envelope that keeps routine fixes out of substantial-modification territory. Then keep the records. AI Sigil links each model in your registry to its monitoring plan, its drift decisions and its evidence, so the answer to an auditor is a report rather than a reconstruction.