Many classifiers output probabilities, and many decisions depend on them being accurate: pricing risk, prioritising reviews, combining predictions. Calibration measures whether those probabilities can be taken at face value.
What Calibration Means
A model is well calibrated if, among all the cases where it predicts 0.8, about 80% are actually positive. A model can rank cases well (high ROC AUC) and still be poorly calibrated.
How to Check
- Reliability diagram: group predictions into bins (0–0.1, 0.1–0.2…) and plot the average prediction against the observed frequency in each bin. A calibrated model follows the diagonal.
- Brier score: the mean squared difference between predicted probabilities and outcomes; lower is better.
Which Models Are Miscalibrated
- Naive Bayes and boosted trees are often overconfident or skewed.
- SVMs don't produce probabilities natively.
- Class weighting and resampling for imbalance distort probabilities.
- Logistic regression is usually reasonably calibrated.
How to Fix It
- Platt scaling: fit a logistic regression on the model's scores.
- Isotonic regression: a flexible, non-decreasing mapping; needs more data.
Fit the calibrator on data the model wasn't trained on, for example with CalibratedClassifierCV.
from sklearn.calibration import CalibratedClassifierCV
calibrated = CalibratedClassifierCV(model, method="isotonic", cv=5).fit(X_train, y_train)
When It Matters
Calibrate whenever probabilities feed a decision threshold, an expected-cost calculation, or a report to people who will read "70%" literally.