Principal component analysis (PCA) is a technique for reducing the number of features while keeping as much information as possible.
The Idea
When features are correlated — height and weight, or many related sensor readings — they carry overlapping information. PCA finds new axes, called principal components, that point in the directions of greatest variation. The first component captures the most variance, the second the most of what remains, and so on. Keeping only the first few components gives a compact summary of the data.
Steps
- Standardise the features (PCA is sensitive to scale).
- Compute the principal components.
- Look at the explained variance ratio to decide how many to keep — often enough to cover 90–95% of variance.
- Transform the data into that smaller set of components.
Uses
- Visualisation: plot high-dimensional data in two dimensions.
- Speed and stability: fewer features for downstream models.
- Noise reduction: dropping low-variance components can remove noise.
- Dealing with multicollinearity in regression.
Limitations
- Components are combinations of original features, so they are harder to interpret.
- PCA captures only linear relationships.
- High variance isn't always what matters for prediction.
Related Methods
For visualising complex structure, t-SNE and UMAP often reveal clusters better than PCA, though their distances are less directly interpretable.
from sklearn.decomposition import PCA
pca = PCA(n_components=0.95).fit(X_scaled)
X_small = pca.transform(X_scaled)