Skip to content

k-Means Clustering

How k-means groups unlabelled data into clusters, how to choose the number of clusters, and its main limitations.

Editorial team 2 min read

k-means is the most widely used clustering algorithm. It groups data into k clusters so that points in a cluster are close to its centre.

How It Works

  1. Choose k and place k initial centres (usually with the smarter k-means++ initialisation).
  2. Assign every point to its nearest centre.
  3. Move each centre to the average of its assigned points.
  4. Repeat steps 2–3 until assignments stop changing.

Choosing k

  • Elbow method: plot the within-cluster sum of squares against k and look for the bend.
  • Silhouette score: measures how well each point fits its cluster versus the next nearest one.
  • Usefulness: often the deciding factor — do the clusters make sense to the people who will use them?

Preparing Data

Scale features first, since k-means uses distances. Remove or cap extreme outliers, which can pull centres away.

Limitations

  • Assumes roughly spherical clusters of similar size.
  • Results depend on initialisation; run it several times (n_init).
  • Every point is forced into a cluster, even outliers.
  • Needs k in advance.

Alternatives

  • DBSCAN finds clusters of arbitrary shape and labels outliers as noise.
  • Hierarchical clustering builds a tree of clusters you can cut at any level.
  • Gaussian mixture models allow soft, probabilistic membership.

Typical Uses

Customer segmentation, grouping documents by topic (using embeddings), image colour quantisation and exploratory analysis.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026