In many real problems the interesting class is rare: fraudulent transactions, faulty parts, positive diagnoses. Imbalance causes models to ignore the minority class.
Why It's a Problem
A model trained to maximise accuracy can simply predict the majority class for everything and look excellent. The rare class — usually the one you care about — gets missed.
Use the Right Metrics
Replace accuracy with precision, recall, F1 and PR AUC for the minority class, and look at the confusion matrix.
Techniques
- Class weights: tell the algorithm to penalise mistakes on the rare class more heavily (
class_weight="balanced"in scikit-learn). Often the simplest effective fix. - Threshold tuning: instead of 0.5, pick the decision threshold that gives the precision–recall balance you need.
- Resampling: undersample the majority class or oversample the minority. SMOTE creates synthetic minority examples; use it with care, and only on training folds.
- Better features: often the real bottleneck is signal, not balance.
- Anomaly detection: when positives are extremely rare or unlabelled, model "normal" behaviour and flag deviations.
Evaluation Pitfalls
- Resample only the training data, never the validation or test data.
- Use stratified splits so every fold contains rare cases.
- Report results at the threshold you will actually use.
Business Framing
Rarity often means each positive is valuable. Estimate the cost of a miss and of a false alarm, and choose thresholds that minimise total cost.