Skip to content

Cohort Analysis

Grouping customers by when they started to understand retention and behaviour over time, with a worked approach.

Editorial team 2 min read

Cohort analysis groups people by a shared starting point — usually the month they first signed up or purchased — and tracks their behaviour over time.

Why Cohorts

Overall metrics mix customers of different ages. Rising total revenue can hide the fact that each new group of customers is less engaged than the last. Cohorts reveal this.

Building a Retention Cohort Table

  1. Assign each customer to a cohort: the month of their first purchase.
  2. For each later activity, compute the period number: months since the cohort month.
  3. Count active customers per cohort and period.
  4. Divide by cohort size to get retention rates.

The result is a triangle: each row a cohort, each column months since start.

first = orders.groupby("customer_id")["order_date"].min().dt.to_period("M").rename("cohort")
orders = orders.join(first, on="customer_id")
orders["period"] = (orders["order_date"].dt.to_period("M") - orders["cohort"]).apply(lambda d: d.n)
table = orders.groupby(["cohort", "period"])["customer_id"].nunique().unstack()
retention = table.div(table[0], axis=0)

Reading the Table

  • Down a column: are newer cohorts retaining better or worse at the same age?
  • Along a row: how does one cohort decay over time?
  • Diagonals: events affecting all cohorts at the same calendar time.

Beyond Retention

Track revenue per customer, average orders or feature adoption by cohort.

Pitfalls

Small cohorts produce noisy rates; recent cohorts have few periods of data; and definitions of "active" must be consistent.

More in Data science & analytics

All Data science & analytics guides →
Data science & analytics Guide · 2 min

Descriptive Statistics Essentials

Mean, median, mode, spread and shape: the summary numbers every analysis starts with, and when each one misleads.

Data science & analytics 2 min read 6 Mar 2026

Data science & analytics Guide · 2 min

Probability Basics for Data Work

The probability ideas analysts use every day: events, conditional probability, independence and Bayes' theorem.

Data science & analytics 2 min read 5 Mar 2026

Data science & analytics Guide · 2 min

Common Probability Distributions

Normal, binomial, Poisson, exponential and more: recognising the shapes data takes and what they imply.

Data science & analytics 2 min read 4 Mar 2026

Data science & analytics Guide · 2 min

Hypothesis Testing Explained

Null hypotheses, p-values and significance: what a hypothesis test tells you, and the misunderstandings to avoid.

Data science & analytics 2 min read 3 Mar 2026