Skip to content

Building Machine Learning Pipelines With scikit-learn

Chain preprocessing and modelling into one object that prevents leakage and deploys cleanly.

Editorial team 2 min read

A scikit-learn Pipeline chains data-preparation steps and a model into a single object. It is the simplest way to avoid data leakage and to ship a model reliably.

Why Pipelines

  • No leakage: each step is fitted only on training folds during cross-validation.
  • Reproducibility: the exact same transformations run at training and prediction time.
  • Deployment: save one object that takes raw inputs and returns predictions.
  • Tuning: search over preprocessing and model hyperparameters together.

Mixed Column Types

A ColumnTransformer applies different steps to different columns:

from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.ensemble import HistGradientBoostingClassifier

prepare = ColumnTransformer([
    ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), numeric_cols),
    ("cat", make_pipeline(SimpleImputer(strategy="most_frequent"),
                          OneHotEncoder(handle_unknown="ignore")), categorical_cols),
])
model = Pipeline([("prepare", prepare), ("clf", HistGradientBoostingClassifier())])
model.fit(X_train, y_train)

Tuning Through a Pipeline

Parameter names use double underscores: clf__learning_rate, prepare__num__simpleimputer__strategy.

Saving and Loading

Save the fitted pipeline with joblib.dump and load it where predictions are needed. Record library versions, because pickled objects can break across versions — and only load files you trust.

Custom Steps

Wrap your own feature logic in a FunctionTransformer or a small class with fit and transform, so it stays inside the pipeline too.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026