Data leakage happens when information that won't be available at prediction time sneaks into training or evaluation. The model looks excellent in testing, then disappoints in production.
Common Forms
- Target leakage: a feature that is a consequence of the outcome. Predicting loan default using "number of collection calls" leaks the answer.
- Train–test contamination: preprocessing (scaling, imputation, feature selection) fitted on the full dataset before splitting.
- Temporal leakage: random splits of time-ordered data let the model learn from the future.
- Group leakage: records from the same person or device on both sides of the split.
- Duplicate leakage: near-identical rows in training and test sets.
Warning Signs
- Performance that seems too good to be true.
- One feature with overwhelming importance.
- A big drop between offline scores and live results.
Prevention Checklist
- For every feature, ask: would I know this at the moment of prediction?
- Split data before any preprocessing, and fit transformations on training data only — use pipelines.
- Split by time for temporal problems and by group for grouped data.
- Remove duplicates across splits.
- Keep the test set locked away until the end.
A Useful Habit
Write down the exact moment a prediction will be made in production, and the data available at that moment. Build the training dataset to mirror it.