Skip to content

Principal Component Analysis

PCA compresses many correlated features into a few components that capture most of the variation. How it works and when to use it.

Editorial team 2 min read

Principal component analysis (PCA) is a technique for reducing the number of features while keeping as much information as possible.

The Idea

When features are correlated — height and weight, or many related sensor readings — they carry overlapping information. PCA finds new axes, called principal components, that point in the directions of greatest variation. The first component captures the most variance, the second the most of what remains, and so on. Keeping only the first few components gives a compact summary of the data.

Steps

  1. Standardise the features (PCA is sensitive to scale).
  2. Compute the principal components.
  3. Look at the explained variance ratio to decide how many to keep — often enough to cover 90–95% of variance.
  4. Transform the data into that smaller set of components.

Uses

  • Visualisation: plot high-dimensional data in two dimensions.
  • Speed and stability: fewer features for downstream models.
  • Noise reduction: dropping low-variance components can remove noise.
  • Dealing with multicollinearity in regression.

Limitations

  • Components are combinations of original features, so they are harder to interpret.
  • PCA captures only linear relationships.
  • High variance isn't always what matters for prediction.

For visualising complex structure, t-SNE and UMAP often reveal clusters better than PCA, though their distances are less directly interpretable.

from sklearn.decomposition import PCA
pca = PCA(n_components=0.95).fit(X_scaled)
X_small = pca.transform(X_scaled)

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026