When this framework fits
Use it when
- Retention, renewal, or repeat use is central to the thesis.
- Reported metrics come from a management dashboard or summary table.
- Mix, customer age, or contract structure may distort the average.
Do not force it when
- The business is project-based and recurrence is not economic retention.
- Raw identifiers cannot be reconciled across systems.
- The observed history is too short for the claimed customer life.
The reconstruction framework
1. Fix the definition before calculating
State the unit, population, clock, and treatment rules. Is the unit a logo, contract, user, location, or dollar of recurring revenue? Does a pause count as churn? Is a win-back treated as continuous retention or a new customer? Are acquisitions, currency effects, and one-time fees excluded? Two analysts can produce different answers from the same ledger while both formulas look correct.
2. Start from the rawest reliable event trail
Use invoices, subscriptions, contracts, product events, or transaction records rather than a precomputed retention export. Preserve stable customer identifiers and a dated record of starts, renewals, upgrades, downgrades, pauses, cancellations, and reactivations. Build a bridge from source records to the analysis population so exclusions can be inspected.
3. Build cohorts and survival curves
Group customers by start period, then follow each cohort through comparable ages. Calendar averages mix mature and immature customers; cohorts reveal whether newer customers retain differently. Logo retention measures the share of customers that remain. Gross revenue retention captures churn and contraction but excludes expansion. Net revenue retention adds expansion. Survival analysis is useful when customers enter at different times and many have not yet had the chance to churn.
4. Segment before you average
Break the result by customer size, product, acquisition channel, contract length, geography, and implementation cohort where sample size permits. A healthy enterprise segment can conceal rapid small-customer churn, while a recent low-quality channel can make a durable core look weak. Weighting matters: one large account can dominate revenue retention while logo behavior deteriorates.
5. Reconcile to the financial statements
The cohort build should explain reported recurring revenue. Create a bridge from opening recurring revenue through new business, expansion, contraction, churn, price, currency, and closing revenue. Differences may be legitimate, but unexplained differences are a control problem and sometimes a thesis problem.
6. Underwrite drivers, not the headline
Model churn hazard by customer age and segment, then connect it to contribution margin, acquisition cost, payback, and cash runway. Use conservative, base, and upside cases. The key sensitivity may not be the long-run retention rate; it may be early-life activation, renewal concentration, onboarding capacity, or a contract cliff.
A reported 95% retention rate becomes 88%
A subscription company reports 95% annual logo retention. Rebuilding monthly cohorts shows that cancelled customers who return within 90 days were counted as continuously retained, several implementation accounts were excluded after launch, and customers on paused contracts remained active. Using a consistent active-contract definition produces 88% annual logo retention.
The average still hides the decision. Enterprise customers retain at 96%, while small customers acquired through one paid channel retain at 71% and require more support. The underwriting case changes from "the product has broad retention" to "the enterprise segment is durable, while the small-customer channel destroys value unless activation and support economics improve."
Evidence requirements
| Evidence | Why it matters | Quality check |
|---|---|---|
| Stable customer and contract IDs | Prevents duplicate or fragmented histories | Resolve mergers and renames explicitly |
| Dated invoices or subscription events | Reconstructs starts, changes, and exits | Reconcile totals to the ledger |
| Product and support events | Tests activation and service burden | Check for missing periods and instrumentation changes |
| Contract terms and renewal dates | Exposes renewal cliffs and cancellation rights | Separate contracted from recognized revenue |
| Acquisition source and customer attributes | Explains mix and channel quality | Suppress segments too small to interpret reliably |
Common failure modes
Definition drift: the metric changes between periods without a restated history. Survivorship bias: only current customers appear in the dataset. Calendar mixing: young cohorts are compared with mature cohorts. Revenue concentration blindness: a few large renewals dominate the average. Win-back inflation: a customer that churned and returned is treated as continuously retained. False precision: a small sample produces a smooth long-run curve that the data cannot support.
Underwriting checklist
- The population, unit, clock, and treatment of pauses and win-backs are documented.
- Raw records include churned customers, not only the current book.
- Logo, gross revenue, and net revenue retention are kept distinct.
- Cohorts are compared at the same age and segmented where decision-relevant.
- The cohort build reconciles to reported recurring revenue with an explicit bridge.
- Concentration, contract cliffs, onboarding, and support burden are tested.
- The downside model uses observed drivers and states data limitations.
Limitations
Historical cohorts may not represent the future after a product, pricing, channel, or macro change. Sparse datasets cannot support fine segmentation, and short histories require assumptions about mature behavior. Retention also does not prove attractive economics if support costs, discounts, or acquisition costs rise. Treat the reconstruction as evidence about the installed system, then make any forward adjustment explicit and challengeable.
Related services and cases
Underwrite the behavior, not the slide
We can rebuild the cohort evidence and connect it to the downside case.