k-means is the most widely used clustering algorithm. It groups data into k clusters so that points in a cluster are close to its centre.
How It Works
- Choose k and place k initial centres (usually with the smarter k-means++ initialisation).
- Assign every point to its nearest centre.
- Move each centre to the average of its assigned points.
- Repeat steps 2–3 until assignments stop changing.
Choosing k
- Elbow method: plot the within-cluster sum of squares against k and look for the bend.
- Silhouette score: measures how well each point fits its cluster versus the next nearest one.
- Usefulness: often the deciding factor — do the clusters make sense to the people who will use them?
Preparing Data
Scale features first, since k-means uses distances. Remove or cap extreme outliers, which can pull centres away.
Limitations
- Assumes roughly spherical clusters of similar size.
- Results depend on initialisation; run it several times (
n_init). - Every point is forced into a cluster, even outliers.
- Needs k in advance.
Alternatives
- DBSCAN finds clusters of arbitrary shape and labels outliers as noise.
- Hierarchical clustering builds a tree of clusters you can cut at any level.
- Gaussian mixture models allow soft, probabilistic membership.
Typical Uses
Customer segmentation, grouping documents by topic (using embeddings), image colour quantisation and exploratory analysis.