A model trained on carefully engineered features will fail if production computes those features differently.
Training–Serving Skew
Skew happens when features at prediction time differ from those in training — different code, different data sources, different timing. It's one of the most common causes of production models underperforming.
Causes
- Features reimplemented separately for training (in a notebook) and serving (in an application).
- Using data at training time that isn't available at prediction time.
- Different handling of missing values or categories.
- Time-zone and date-boundary differences.
- Stale features in production.
Preventing It
- Define features once in shared code used by both training and serving, or in a feature store.
- Point-in-time correct training data: use feature values as they were when each training example occurred.
- Log features at prediction time and compare their distributions with training data.
- Test that the same input produces the same features in both paths.
Freshness
Decide how fresh each feature must be — recomputed daily, hourly or on each request — and monitor that it's actually updated.
Keep Features Simple
Complex features with many dependencies are harder to compute reliably in production. Prefer features that are robust, well understood and available on time.
Document Features
Record the definition, source, owner and freshness of each feature so others can reuse and maintain them.