Skip to content

Anomaly Detection

Finding unusual records — fraud, faults, intrusions — with statistics and machine learning, even when examples of anomalies are scarce.

Editorial team 2 min read

Anomaly detection finds observations that differ markedly from the norm. It's used for fraud detection, equipment monitoring, cybersecurity and data quality checks.

Types of Anomalies

  • Point anomalies: a single unusual value, such as a huge transaction.
  • Contextual anomalies: normal in one context but not another — high electricity use at 3am.
  • Collective anomalies: a sequence that is unusual as a whole.

Approaches

  • Statistical rules: z-scores, interquartile-range limits, control charts. Simple and explainable.
  • Isolation Forest: isolates points with random splits; anomalies are isolated quickly.
  • Local Outlier Factor: compares a point's density with its neighbours'.
  • One-class SVM: learns a boundary around normal data.
  • Autoencoders: neural networks that reconstruct normal data well and anomalies poorly.
  • Supervised classifiers: when labelled anomalies exist, standard classification (with class imbalance handling) often works best.

The Challenge of Evaluation

Labelled anomalies are usually rare or missing. Use whatever confirmed cases exist, have experts review top-ranked alerts, and track precision on reviewed cases over time.

Practical Tips

  • Define "normal" carefully; include seasonality and context.
  • Tune the alert threshold to reviewers' capacity.
  • Expect drift: what's normal changes, so retrain regularly.
  • Combine anomaly scores with business rules to reduce false alarms.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026