Cohort analysis groups people by a shared starting point — usually the month they first signed up or purchased — and tracks their behaviour over time.
Why Cohorts
Overall metrics mix customers of different ages. Rising total revenue can hide the fact that each new group of customers is less engaged than the last. Cohorts reveal this.
Building a Retention Cohort Table
- Assign each customer to a cohort: the month of their first purchase.
- For each later activity, compute the period number: months since the cohort month.
- Count active customers per cohort and period.
- Divide by cohort size to get retention rates.
The result is a triangle: each row a cohort, each column months since start.
first = orders.groupby("customer_id")["order_date"].min().dt.to_period("M").rename("cohort")
orders = orders.join(first, on="customer_id")
orders["period"] = (orders["order_date"].dt.to_period("M") - orders["cohort"]).apply(lambda d: d.n)
table = orders.groupby(["cohort", "period"])["customer_id"].nunique().unstack()
retention = table.div(table[0], axis=0)
Reading the Table
- Down a column: are newer cohorts retaining better or worse at the same age?
- Along a row: how does one cohort decay over time?
- Diagonals: events affecting all cohorts at the same calendar time.
Beyond Retention
Track revenue per customer, average orders or feature adoption by cohort.
Pitfalls
Small cohorts produce noisy rates; recent cohorts have few periods of data; and definitions of "active" must be consistent.