Glossary
Concept Drift
Concept drift is a change in the statistical relationship between a model's inputs and the target label over time, so the mapping learned during training no longer holds. As the conditional distribution P(y|X) shifts, prediction accuracy degrades even when the input features themselves look unchanged.
How Concept Drift Works
Concept drift works through a change in the conditional probability P(y|X), the likelihood of a label given the input, rather than the inputs alone. Gama et al. (2014), in the ACM Computing Surveys paper "A Survey on Concept Drift Adaptation," classify drift by speed and pattern: sudden drift replaces one relationship abruptly, gradual drift blends old and new over an interval, and recurring drift reintroduces a past relationship, often seasonally.
A model learns a fixed function during training. It maps an input X, the pixels and tokens of a document, to a label y, the field value a reader wants. Concept drift means the correct y for a given X has moved since the model was fit. The model keeps applying the old function and returns wrong answers with unchanged confidence.
Detection typically compares a stream of recent predictions against a baseline. Because the raw inputs can look stable while the label relationship moves, error-based detectors watch accuracy or a proxy for it and flag a sustained rise. Once a detector confirms drift rather than a one-off outlier, the standard response is retraining on freshly labeled data, so the function is refit to the current relationship.
Why Concept Drift Matters
Concept drift matters because a commercial real estate extraction model decays silently as document formats and market terms change. A model trained to read last year's lease abstracts and offering memoranda keeps returning confident answers, but the relationship between the text it sees and the correct field value has moved, so accuracy falls without any error message.
The failure is quiet, which is the risk. An underwriter pulling net operating income, rent escalations, or reimbursement terms from an offering memorandum sees a populated field, not a warning. If a lender restructures how it labels a debt yield covenant, or a broker adopts a new template for a rent roll, the model can extract the old pattern into the new field and feed a wrong number into the deal model. On a mispriced acquisition that error compounds through every downstream calculation.
Example
The example is a rent-roll and lease extraction model watched across a market cycle. Each drift type below shows how the input-to-label relationship shifts in commercial real estate documents and the signal that flags it. Detection compares recent labeled predictions against a baseline, since the raw inputs can look stable while accuracy erodes.
Drift type | CRE example | Detection signal |
|---|---|---|
Sudden | A lender switches its loan-agreement template overnight, moving the maturity date to a new section | Extraction accuracy on maturity fields drops sharply in one week |
Gradual | Brokers phase in a new offering-memorandum format, so old and new layouts circulate together for a quarter | Error rate climbs steadily as the new layout's share grows |
Recurring | Year-end reforecasts reintroduce a reporting style seen every fourth quarter | Accuracy dips on a predictable seasonal cycle |
Incremental | Market vocabulary shifts as "expense stop" gives way to "base year" phrasing over many deals | Small, sustained rise in reimbursement-clause errors |
Variations and Edge Cases
Variations of concept drift differ by speed, recurrence, and whether the change is real or virtual. Real drift moves P(y|X), the true relationship. Virtual drift moves the input distribution P(X) without changing the correct label. Incremental drift shifts through many small steps, while blips are one-off outliers that should not trigger retraining.
The edge case that wastes the most effort is treating a blip as drift. A single malformed scan or an unusual one-off document raises the error rate for a moment, but the underlying relationship has not changed. Retraining on it teaches noise. Well-tuned detectors require a sustained shift, not a single spike, before they signal. The opposite failure is a slow incremental drift that never crosses a threshold on any single day yet moves accuracy materially over a quarter.
Concept Drift vs Data Drift
Concept drift is often confused with data drift. Concept drift is a change in the relationship P(y|X) between inputs and the label, so the correct answer for the same input changes. Data drift, also called covariate shift, is a change in the input distribution P(X) alone, while the input-to-label mapping stays fixed. Both degrade accuracy.
Dimension | Concept drift | Data drift |
|---|---|---|
What changes | P(y | X), input-to-label relationship |
CRE example | A covenant term is redefined, so the same clause maps to a new value | A new asset class enters the pipeline with unfamiliar layouts |
Fix | Retrain on freshly labeled data | Broaden training coverage; relabel may be unnecessary |
The practical difference: data drift can sometimes be handled by exposing the model to more of the new inputs, while concept drift requires new labels because the definition of the right answer has moved.
Frequently Asked Questions
What causes concept drift? Concept drift is caused by any change in the real-world process that links inputs to outcomes: new document templates, revised contract language, shifting market conventions, or regulatory changes. The inputs may look similar while the correct label for them moves.
How is concept drift detected? Concept drift is detected by monitoring a model's error rate or a proxy for it against a baseline and flagging a sustained increase. Error-based detectors from the concept-drift literature confirm a real shift rather than reacting to a single outlier.
How do you fix concept drift? The standard fix is retraining the model on freshly labeled data that reflects the current relationship. Recurring drift can be handled by keeping and reusing past models when a known pattern returns, avoiding retraining from scratch.