Skip to content

Handling Imbalanced Classes

When one class is rare — fraud, defects, disease — standard training and metrics mislead. Techniques that help.

Editorial team 2 min read

In many real problems the interesting class is rare: fraudulent transactions, faulty parts, positive diagnoses. Imbalance causes models to ignore the minority class.

Why It's a Problem

A model trained to maximise accuracy can simply predict the majority class for everything and look excellent. The rare class — usually the one you care about — gets missed.

Use the Right Metrics

Replace accuracy with precision, recall, F1 and PR AUC for the minority class, and look at the confusion matrix.

Techniques

  • Class weights: tell the algorithm to penalise mistakes on the rare class more heavily (class_weight="balanced" in scikit-learn). Often the simplest effective fix.
  • Threshold tuning: instead of 0.5, pick the decision threshold that gives the precision–recall balance you need.
  • Resampling: undersample the majority class or oversample the minority. SMOTE creates synthetic minority examples; use it with care, and only on training folds.
  • Better features: often the real bottleneck is signal, not balance.
  • Anomaly detection: when positives are extremely rare or unlabelled, model "normal" behaviour and flag deviations.

Evaluation Pitfalls

  • Resample only the training data, never the validation or test data.
  • Use stratified splits so every fold contains rare cases.
  • Report results at the threshold you will actually use.

Business Framing

Rarity often means each positive is valuable. Estimate the cost of a miss and of a false alarm, and choose thresholds that minimise total cost.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026