Features often come in wildly different units — age in years, income in dollars, ratios between zero and one. Some algorithms care; others don't.
Which Algorithms Need Scaling
- Distance-based: k-nearest neighbours, k-means, SVMs.
- Gradient-based: linear and logistic regression with regularisation, neural networks.
- PCA, which finds directions of maximum variance.
Tree-based models — decision trees, random forests, gradient boosting — split on thresholds and are unaffected by scaling.
Common Methods
- Standardisation (z-score): subtract the mean and divide by the standard deviation. Features end up centred at zero with unit variance. A good default.
- Min-max scaling: rescale to a fixed range, usually 0 to 1. Sensitive to outliers.
- Robust scaling: use the median and interquartile range, reducing outlier influence.
- Log transform: compress long-tailed features such as income or counts before scaling.
Avoid Leakage
Fit the scaler on the training data only, then apply the same transformation to validation and test data. Scaling the whole dataset first leaks information from test into training. Pipelines make this automatic.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_train, y_train)
Don't Forget Production
The fitted scaler is part of the model. Save it with the model so new data is transformed exactly as training data was.