Skip to content

Feature Engineering Pipelines in Production

Keeping features consistent between training and serving, computing them on time, and avoiding training–serving skew.

Editorial team 2 min read

A model trained on carefully engineered features will fail if production computes those features differently.

Training–Serving Skew

Skew happens when features at prediction time differ from those in training — different code, different data sources, different timing. It's one of the most common causes of production models underperforming.

Causes

  • Features reimplemented separately for training (in a notebook) and serving (in an application).
  • Using data at training time that isn't available at prediction time.
  • Different handling of missing values or categories.
  • Time-zone and date-boundary differences.
  • Stale features in production.

Preventing It

  • Define features once in shared code used by both training and serving, or in a feature store.
  • Point-in-time correct training data: use feature values as they were when each training example occurred.
  • Log features at prediction time and compare their distributions with training data.
  • Test that the same input produces the same features in both paths.

Freshness

Decide how fresh each feature must be — recomputed daily, hourly or on each request — and monitor that it's actually updated.

Keep Features Simple

Complex features with many dependencies are harder to compute reliably in production. Prefer features that are robust, well understood and available on time.

Document Features

Record the definition, source, owner and freshness of each feature so others can reuse and maintain them.

More in MLOps & deployment

All MLOps & deployment guides →
MLOps & deployment Guide · 2 min

What Is MLOps?

The practices that take machine learning from notebook to reliable production: versioning, automation, deployment and monitoring.

MLOps & deployment 2 min read 26 Dec 2025

MLOps & deployment Guide · 2 min

Deploying Machine Learning Models

Batch scoring, real-time APIs, streaming and on-device inference: choosing how predictions reach users, and deploying safely.

MLOps & deployment 2 min read 25 Dec 2025

MLOps & deployment Guide · 2 min

Monitoring Machine Learning Models in Production

What to monitor after deployment — data, predictions, outcomes and operations — and how to respond when things change.

MLOps & deployment 2 min read 24 Dec 2025

MLOps & deployment Guide · 2 min

Model Registries and Versioning

Why every production model needs a version, metadata and lineage, and how a model registry manages promotion and rollback.

MLOps & deployment 2 min read 23 Dec 2025