Wise Owl Collective white paper
Using Customer Data to Improve Retention and Value
How to choose useful actions and measure the value they create
Executive summary
Customer data becomes valuable when it helps a business choose an action that improves the customer relationship and produces an economic return. Knowing who may leave, who spends the most, or who is likely to respond can inform that decision. None of those predictions, by itself, establishes that a particular intervention will help.
We recommend building retention programs around incremental customer value: the contribution an action creates beyond what would have happened without it, after the cost of the action and its consequences. This changes the question from who looks valuable or at risk to where a useful intervention is likely to make a worthwhile difference.
This paper is for marketing, customer, and analytics leaders planning retention, replenishment, onboarding, and win-back programs. It explains how to connect customer data to an action, test that action fairly, and evaluate the result in economic terms. The approach applies to both recurring contracts and repeat-purchase businesses, with different definitions of retention and different observation windows.
What a useful program must establish
A useful program needs an eligible audience, a reason for the intervention, a deliverable experience, and a comparison that can reveal its effect. A service recovery call, replenishment reminder, education sequence, or offer should address a plausible customer need. Data can help identify that need, but the response still depends on the quality and timing of the experience.
A campaign can raise response while reducing profit if it discounts purchases that would have happened anyway. A churn model can be accurate while targeting customers who will not benefit from the available action. A short-term order increase can pull future demand forward. The measurement plan must distinguish these possibilities before a team treats a campaign result as evidence of lasting value.
Separate customer prediction from action selection
Three questions often become compressed into one score: who is likely to leave, what a continuing relationship may be worth, and who will respond differently because of an intervention. They require different evidence. Keeping them separate makes targeting decisions easier to explain and evaluate.
| Question | What the estimate describes | What it does not establish |
|---|---|---|
| Churn risk | Likelihood of a defined lapse or cancellation within a stated period. | Whether a reminder, call, or offer will change the outcome. |
| Customer value | Expected future contribution under stated assumptions. | The additional value created by a specific action. |
| Incremental effect | The difference an action is expected to make relative to an alternative. | A guaranteed outcome for an individual customer. |
Eva Ascarza’s research combined two field experiments with machine-learning methods and found that customers with the highest predicted churn risk were not necessarily the best targets for proactive retention. The study supports a practical distinction: sensitivity to an intervention matters alongside the risk of departure. This does not justify excluding every high-risk customer. [1]
Begin with an actionable hypothesis
Describe the customer problem and the proposed response. For example, recent buyers may fail to reorder because they do not know when a product needs replenishment. A timely reminder is a testable intervention. If the barrier is product dissatisfaction or a delivery failure, a discount may be a poor substitute for resolving it.
Use a simple targeting rule as an initial comparator. It may be based on time since purchase, a missed onboarding step, or a documented service issue. Introduce more complex modeling when the data and operating scale justify it, and require the model to improve on the simpler approach at the actual budget and contact capacity.
Build data around the decision
Start with the information needed to determine eligibility, deliver the action, and observe its outcome. A retention experiment rarely requires a complete reconstruction of the customer data estate. It does require dependable identifiers, consistent time definitions, and a way to connect assignment and outcomes without silently losing records.
| Data element | Purpose | Check before use |
|---|---|---|
| Customer or account identifier | Connect history, assignment, delivery, and outcomes. | Duplicates, household overlap, changed identifiers. |
| Transaction or contract history | Define the baseline and the outcome window. | Returns, cancellations, missing channels, time zones. |
| Action and delivery records | Distinguish assignment from attempted and actual delivery. | Suppression, failed delivery, competing contacts. |
| Contribution and action cost | Assess value after product and campaign costs. | Margin definition, discounts, service and fulfillment costs. |
| Preferences and eligibility rules | Apply the intended contact and data-use boundaries. | Current preferences, exclusions, access restrictions. |
Define retention for the business you operate. A subscription cancellation has an observable event date. A customer who buys intermittently may simply be between purchases. In the latter setting, define a purchase window and examine the normal buying cadence before treating inactivity as departure. Report repeat purchase, frequency, and contribution separately when a single retention label would obscure the behavior.
Freeze the information available at the moment of selection. An open, click, redemption, or purchase that occurs after assignment must not become a selection feature for the same test. Distinguish an actual zero from a missing outcome, and make sure late returns or cancellations are treated consistently in both groups.
Use only information appropriate to the purpose and permitted contact process. Limit access to sensitive records, retain a record of exclusions, and review whether attributes or proxies lead to unfair or inappropriate treatment. A commercially attractive targeting rule still needs an acceptable customer experience.
Design a comparison that can answer the question
Where feasible, randomly assign eligible customers to the proposed action or the existing approach. Define that existing approach precisely. It may mean normal service and ordinary communications without the incremental campaign, rather than no contact of any kind. Protect necessary service communications and apply the same eligibility rules to both groups.
Choose the assignment unit to reflect how the experience spreads. Customers sharing an account, household, or sales representative may influence one another or receive overlapping treatment. Assigning at an appropriate group level can reduce contamination, but it changes the sample-size and analysis requirements. Document the choice before launch.
| Before launch | Agree and record |
|---|---|
| Primary outcome | A contribution, purchase, or retention measure over a fixed horizon. |
| Commercial threshold | The smallest improvement worth acting on after cost. |
| Customer limits | Acceptable complaints, opt-outs, returns, or service burden. |
| Test design | Eligibility, assignment unit, allocation, sample size, and analysis plan. |
| Observation period | Enough time for the expected response and relevant downstream effects. |
Analyze the effect of assignment using the planned population. Restricting analysis to people who opened or redeemed can destroy the original comparability because those behaviors occur after assignment. Delivery and engagement measures are valuable diagnostics, but they answer different questions from the effect of offering the program.
Verify that group counts, identifiers, and outcome capture behave as intended. Microsoft’s experimentation research highlights sample ratio mismatch as a warning about assignment, execution, or analysis integrity. Investigate the cause rather than adjusting the data until the groups look convenient. [2]
Plan sample size using the baseline rate, the smallest commercially meaningful effect, and the uncertainty the decision can tolerate. Avoid repeated unplanned checks followed by stopping at a favorable result. If the design must change, record what changed and how that affects interpretation. When a randomized test is impossible, describe the assumptions behind the alternative and the sources of bias it cannot remove.
Measure contribution after the cost of the action
The following hypothetical example shows why an increase in repeat purchase is not enough to justify a campaign. Suppose 20,000 eligible customers are split equally between a control group and an offer group. In the defined period, each purchasing customer makes one qualifying order. Contribution is $40 per order before the campaign discount and contact cost, after ordinary product and variable fulfillment costs.
| Illustrative result | Control | Offer |
|---|---|---|
| Assigned customers | 10,000 | 10,000 |
| Customers purchasing | 1,500 | 1,700 |
| Purchase rate | 15% | 17% |
| Contribution before campaign costs | $60,000 | $68,000 |
| Discount on every offer-group order | $0 | $17,000 |
| Incremental contact cost | $0 | $1,000 |
| Contribution after campaign costs | $60,000 | $50,000 |
The observed purchase-rate difference is 2 percentage points, or a 13.3% relative lift. Applied to the 10,000 customers in the offer group, that difference corresponds to 200 additional purchasing customers and $8,000 of additional contribution before campaign costs. But a $10 discount on all 1,700 offer-group orders costs $17,000. Contact at $0.10 per assigned customer adds $1,000. The observed contribution difference is therefore a $10,000 loss.
These are illustrative assumptions, not client results. A real analysis should estimate uncertainty for the economic outcome, account for multiple orders and returns, and observe whether the action changes later behavior. A positive response effect may still be commercially unattractive. A negative short-term contribution result should not be excused by an unsupported lifetime-value forecast.
Keep attribution and incrementality distinct
Attribution connects recorded conversions to touchpoints under a chosen rule. It does not establish which purchases would have occurred without the campaign. Microsoft’s incrementality guidance makes this distinction explicitly. Use the experimental comparison to estimate the additional effect and use channel diagnostics to understand how the experience operated. [3]
Improve targeting after establishing a credible effect
Once a program produces reliable experimental data, examine whether its effect varies across customer groups. Uplift modeling, also called treatment-effect modeling in related settings, attempts to estimate how the effect of an action varies with observed characteristics. A survey by Zhang, Li, and Liu connects these approaches under a common causal framework. Such estimates concern patterns across comparable customers; an individual’s unobserved alternative outcome remains unavailable. [4]
Use separate data to develop and evaluate the targeting policy. A segment selected because it performed well in one sample may not repeat that result. Compare the proposed policy with broad eligibility and simple rules at the same contact capacity and budget. Report expected contribution, uncertainty, and the amount of the eligible audience for which the evidence is weak.
Connect selection to the delivered experience
A useful model still needs current data, an available offer or service action, and a channel that can deliver it. Specify contact limits, suppression rules, escalation routes, and coordination with other campaigns. Log what was selected, what was sent, and what the customer could actually receive. A recommended action that arrives too late or conflicts with another message is a different intervention from the one tested.
Keep a suitable comparison as the program expands, where feasible, and review outcomes by cohort and period. Watch for changes in product availability, seasonality, customer mix, and the costs of delivering the action. Investigate complaints, opt-outs, repeated discount use, and movement between channels. A program can shift when and where purchases occur without increasing total value.
Set a limit on lifetime value assumptions
A customer-value forecast is a model of future contribution under assumptions about purchasing, retention, cost, and time. Its usefulness depends on how those assumptions hold up. Begin with a horizon the business can observe, then extend the view as evidence accumulates. Compare forecasts with realized cohort outcomes and show how sensitive the decision is to the longer-term assumptions.
Scale gradually when the estimated benefit is economically meaningful, customer outcomes remain acceptable, and the operating process is reliable. If evidence is weak, simplify the action, narrow the population, or extend observation. More elaborate targeting cannot repair an intervention that does not solve a useful customer problem.
Prepare the first retention experiment
A concise experiment brief can align marketing, operations, finance, and analytics before a program reaches customers. It should make the following decisions explicit.
- Customer need: What problem or friction is the action intended to address?
- Eligible audience: Who qualifies, who is excluded, and what information is available at selection?
- Action and comparison: What experience changes, and what will the comparison group receive?
- Outcome and horizon: What will be measured, over what period, and at what unit of assignment?
- Economics and customer limits: What improvement is worth the cost, and what adverse outcomes require action?
- Ownership and next step: Who can change delivery, investigate data issues, and approve any expansion?
This brief should be specific enough to expose disagreements before launch. If finance defines value as contribution and marketing defines success as response, reconcile the measures first. If service teams cannot deliver the proposed intervention consistently, fix the delivery process before testing its effectiveness.
Work with Wise Owl Collective
Wise Owl Collective helps organizations connect customer insight to practical marketing, loyalty, and retention decisions. Bring a customer lifecycle problem, the action you are considering, and the data you can observe. We can help define an appropriate comparison, connect delivery to measurement, and assess whether the result supports broader use.
Discuss your retention opportunity
References
- Eva Ascarza. Retention Futility Targeting High-Risk Customers Might Be Ineffective. Journal of Marketing Research 55(1), 80-98, 2018.
- Microsoft Research. Diagnosing Sample Ratio Mismatch in A/B Testing. September 2020.
- Microsoft Learn. Data Science Toolkit Incrementality. Updated October 2025.
- Weijia Zhang, Jiuyong Li, and Lin Liu. A Unified Survey of Treatment Effect Heterogeneity Modeling and Uplift Modeling. 2020.
Sources reviewed October 2026. The worked example is hypothetical and demonstrates the economics of a test; it is not a performance claim or forecast.