Skip to content
Subconscious

Why a 20% Retention Rate Means Different Things for a 10-User Cohort and a Million-User Cohort

A marketing analytics lead sees 20 percent retention. Two active users out of ten and 200,000 out of a million have the same observed ratio, but different sampling uncertainty under a binomial model. Report the numerator, denominator, measurement window and assumptions before using the rate to compare acquisition channels.

A retention ratio needs its denominator

Retention is always between zero and one, and it's a quotient: active users divided by cohort size. Two cohorts can post the same 20% and carry completely different amounts of information.

CohortCohort sizeRetention rateHow much that number tells you
A10, with 2 active20%Wilson 95% interval approximately 5.7% to 51.0%
B1,000,000, with 200,000 active20%Wilson 95% interval approximately 19.92% to 20.08%

These are invented cohorts and conditional binomial intervals. A dashboard can use a Wilson proportion interval from the count and denominator without building a new forecasting model. A binomial regression can also represent rates with their trial counts.

How do you model active users instead of a ratio?

A binomial likelihood ties the number of active users to cohort size and an underlying retention probability:

N_active ~ Binomial(N, p)

Here N is cohort size and p is the retention probability. A binomial model assumes appropriately measured, conditionally independent trials with a common probability within the modeled group. Shared acquisition or operational shocks, repeated users and selection can violate that simplification. Consider cohort/time effects or overdispersion, and validate prediction intervals on future cohorts.

A companion revenue model follows the same logic: revenue per cohort-period is modeled with a distribution tied to the number of active users and an average-revenue-per-user parameter, so retention and revenue share the same cohort-level uncertainty rather than being reconciled after the fact.

What this buys forecasting, specifically

Juan Orduz's Cohort Revenue and Retention preprint and PyMC Labs tutorial illustrate a joint forecasting approach using binomial retention and BART. Such a model can borrow information across cohorts and link revenue uncertainty with activity; its usefulness depends on predictive checks and held-out performance.

Where does this discipline generalize?

Report uncertainty together with assumptions and external validation. An interval alone does not show that a channel caused retention or that a forecast transfers to a new cohort.

What are the limitations, and what doesn't this prove?

This is a modeling-methods explainer, demonstrated on a synthetic dataset with known ground truth so the model's accuracy could be checked against it. It's a useful pattern for anyone building cohort-level CLV forecasts, not a validated production benchmark, and it doesn't establish results in a live, messy revenue dataset. Extending it, by layering in acquisition channel as a covariate, pooling across hierarchical markets, or swapping BART for a neural-network component, is a reasonable next step for a team with the engineering capacity to build and maintain it, not a guarantee of the same interval widths or forecast accuracy on a different business.

Next step

For a young cohort, report counts, window, interval construction and shared-shock assumptions using the proportion-interval reference above. A claim that a channel caused retention needs a separate identification design. The public research record reports aggregate choice-parameter-rank evidence; request a scoped interval-reporting example and matched cohort outcome evidence for the proposed study. See the study workflow when defining that brief.

Five retention checks: report active count and cohort size, compute a proportion interval, check shared shocks and measurement, use a joint model when forecasting requires it, and validate future cohorts.
A denominator-aware interval can describe uncertainty without a new joint forecasting model.