PlatformDTC
EnterprisePricingAbout UsAnswersBlogDocs
  1. Home/
  2. Glossary/
  3. Cohort

Measurement

Cohort

A cohort is a group of customers who share a starting characteristic — most often the period in which they made their first purchase — and are then tracked together over time.

Cohorting is what separates a change in the business from a change in the mix. A blended metric moves when new customers behave differently, when old customers behave differently, or when their proportions shift, and it cannot tell you which.

Cut cohorts by acquisition channel as well as date. Customers acquired on a heavy first-order discount retain differently, and blending them with organic ones makes both curves unreadable.

The classic misdiagnosis it prevents: overall retention looks flat, so the product team concludes nothing changed. Cohorting reveals that each new cohort is in fact retaining better while the mix has shifted toward a cheaper acquisition channel that retains worse — two large opposing movements averaging to a straight line. Acting on the flat number means fixing a product that is improving and scaling a channel that is not.

What defines a cohort, and why acquisition month is the default

A cohort needs two things: a shared starting event, and a clock that starts at that event. The second is the part people skip and the part that does all the work. Once every customer's timeline is re-indexed so that month 0 is their own first month rather than a calendar month, customers acquired in January and in July become directly comparable at the same age. Without that re-indexing you have a filter, not a cohort.

Acquisition month is the default starting event for a simple reason: it is the one event every customer has exactly once, it never changes retrospectively, and it is recorded reliably. Contrast that with cohorting on "first subscription" — a customer can subscribe, cancel and subscribe again, so the starting event is ambiguous and cohort membership can move after the fact, which quietly rewrites history in every chart built on it.

Acquisition date is not the only useful axis, only the safest one to start from. Behavioural cohorts — first product purchased, acquisition channel, full price versus first-order discount, device, country — are frequently more diagnostic, and they combine: the March cohort acquired on paid social at 20% off is a legitimate cohort and often the one that explains a retention problem. The rule that survives all of these cuts is that the time index must be age, not calendar date.

Monthly is the standard grain for ecommerce because most repeat cycles are measured in weeks to months. Weekly cohorts for a business doing a few hundred orders a month are mostly noise; quarterly cohorts for a business doing tens of thousands hide too much. Choose the grain that gives you cohorts large enough to read, then leave it alone.

Reading a cohort triangle

The standard output is a triangle: one row per cohort, one column per month of age, and progressively fewer cells as you move down to cohorts that have not existed long enough to fill them. The shape is not a defect. It is the honest representation of the fact that recent cohorts have less history, and any chart that hides it is averaging over data that does not exist.

There are three ways to read it and they answer three different questions. Across a row is the ageing of one cohort: how a single group of customers decays. Down a column is the comparison that actually matters for decisions — are newer cohorts better than older ones at the same age? Along the diagonal is one calendar month experienced by every cohort simultaneously, which is where site outages, stockouts and promotions show up as a stripe cutting across cohorts of every age.

Acquisition cohortCustomersMonth 1Month 2Month 3Month 4
January1,20022%15%12%10%
February1,35023%16%13%—
March1,90025%17%——
April2,40026%———
Example cohort triangle: repeat-purchase rate by months since first order

The mistakes that make a cohort chart lie

Read the month 1 column in that example: 22%, 23%, 25%, 26%. Newer cohorts are repeating sooner, which is a real improvement. Now notice that the cohorts are also growing — 1,200 to 2,400 — so the blended repeat rate across all customers is being pulled toward the behaviour of the newest and largest group, and could move in almost any direction depending on how the older cohorts age. The cohort view and the blended view can point in opposite directions without either being wrong.

  • Comparing cohorts of unequal age. April has had one month and January has had four, so any cumulative measure favours January by construction. Always compare at equal age — month 1 against month 1 — and never rank cohorts by a lifetime-to-date figure.
  • Survivorship in the denominator. A retention curve computed over "customers still active" measures survivors, not the cohort. The denominator has to stay the original cohort size for the entire life of the curve, or the number rises as the least engaged customers drop out of the calculation.
  • Right-censoring at the bottom. The newest cohorts have incomplete rows. Averaging them into a summary drags it down; excluding them entirely biases it up by removing the most recent evidence. Show the triangle and let the gaps be visible.
  • Cohorting on a repeatable event. First subscription, first app install, first login — any of these can happen twice, which makes cohort membership mutable and every historical chart unstable.
  • Calendar effects landing at different ages. Black Friday is month 1 for the October cohort and month 11 for the previous December's, so a seasonal spike appears at a different position on every curve. Averaging the curves smears one real event across the whole horizontal axis.
  • Changing the cohort definition without restating history. A cohort table where "new customer" meant something different before June is two charts drawn on one axis.

When the numbers are too small to act on

Cohort analysis fails quietly on small books because a percentage of a small number moves in large steps. A cohort of 40 customers with a 20% month-1 repeat rate is 8 people; two customers either way is 15% or 25%, and one delayed shipment can produce that. A chart of such cohorts looks like a business with wildly volatile retention, and it is really a business with 40-customer cohorts.

The practical response is to widen the grain rather than to smooth the chart — quarterly cohorts of 500 are readable where monthly cohorts of 160 are not — and to put cohort size in a column next to the rates so nobody reads a percentage without knowing what it is a percentage of. The example triangle above does this deliberately.

Then act on the column comparison rather than the row. The row tells you customers decay, which they always do. The column tells you whether the cohorts you are acquiring this quarter are better at the same age than the ones you acquired last quarter, and that is the only question in the chart that anyone can do anything about.

Frequently asked questions

What is a cohort?
A cohort is a group of customers who share a starting characteristic — usually the month of their first purchase — and are tracked together over time, with each customer's clock starting at their own first purchase rather than at a calendar date. That re-indexing by age is what makes groups acquired at different times comparable.
What is cohort analysis?
Cohort analysis groups customers by when they were acquired and follows each group forward, so you can see whether behaviour changed or whether the mix of customers changed. A blended metric moves for either reason and cannot distinguish them, which is how flat overall retention can hide two large opposing trends.
Why cohort by acquisition month?
Because first purchase is the one event every customer has exactly once, it is recorded reliably, and it never changes retrospectively. Cohorting on a repeatable event such as first subscription makes membership ambiguous, since a customer can cancel and start again, and that quietly rewrites every historical chart built on it.
How do you compare two cohorts fairly?
Compare them at the same age, not at the same calendar date. A cohort acquired four months ago has four months of history and one acquired last month has one, so any cumulative figure favours the older one by construction. Read down a column of the cohort table — month 1 against month 1 — rather than across rows.
What is survivorship bias in cohort analysis?
It occurs when the denominator of a retention curve is customers still active rather than the original cohort size. Measured that way, the rate rises as the least engaged customers drop out of the calculation, and a deteriorating cohort can look like an improving one. The denominator must stay fixed at the starting cohort size throughout.
How large does a cohort need to be?
Large enough that ordinary variation does not swamp the signal. A 40-customer cohort with a 20% repeat rate is eight people, so two customers either way swings it between 15% and 25%. Widen the grain to quarterly cohorts rather than smoothing the chart, and always display cohort size next to the percentage.

Related terms

  • Cohort retention
  • RFM (recency, frequency, monetary)
  • LTV (customer lifetime value)
  • Churn rate
  • CAC (customer acquisition cost)
  • Payback period

One platform for the whole order lifecycle

Storefronts, subscriptions, payments, inventory and fulfilment on one system — operated by agents through a scoped, audited gateway.

Talk to salesCheck your store — free

PlatformDTC

One platform to run your brand. Agents included.

Resources

  • Answers
  • Glossary
  • Agent Readiness Checker
  • DTC AI Crawler Index
  • Blog
  • Pricing
  • Explore all pages

Company

  • About Us
  • Enterprise
  • Talk to Sales
  • Contact
  • Developer Docs
  • System Status
  • Community

Legal

  • Terms of Service
  • Privacy Policy
  • Security
  • All policies

© Copyright 2026 PlatformDTC. All Rights Reserved.