Measuring customer experience with control groups: how to hold out
Control groups in customer experience measurement are the honest way to show a change caused a result. How to hold out without A/B tools, and what breaks it.
Table of contents
- Key takeaways
- What a control group is in customer experience measurement
- Why before-and-after comparisons are a story, not evidence
- Five ways to hold out when the A/B tools cannot reach
- How to run a minimum viable control group
- What quietly breaks a control group
- When a control group is not the answer
- Where to start
- FAQ
The new complaint-handling process went live in January. By April, churn among customers who had contacted support was down. The CX team put the two facts on one slide, and for about a week it was the best slide in the company. Then someone from finance pointed out that a price increase had been withdrawn in February, and someone from sales mentioned that the main competitor had spent March apologizing for an outage. The slide was still true. It just no longer proved anything, because the customer experience measurement had no control group.
A control group is a set of customers, chosen to be comparable to the rest, who deliberately do not receive a change, so that whatever happens to them shows what would have happened anyway. The difference between that group and the customers who did get the change is the part the change can claim.
That is the position most experience teams are in when they try to claim a result: they know what changed, they know what happened afterward, and they have nothing that connects the two. The only honest way to connect them is to compare against customers who did not get the change.
Key takeaways
- A before-and-after comparison shows what happened over time but cannot say why, because everything else changed over the same time.
- A control group holds time constant: two comparable groups live through the same quarter and only one gets the change, so the difference between them is the effect.
- You do not need a testing platform; a branch, a region, a cohort, a staggered rollout or a random slice of customers can each serve as a holdout.
- The four things that ruin a control group are contamination, groups too small to show the effect, letting the treatment group choose itself, and hunting for whichever metric moved.
- Never hold out safety fixes, refunds or corrections of genuine errors; stagger instead of withholding wherever you can, and keep the holdout brief.
What a control group is in customer experience measurement
In marketing, a holdout on a campaign is routine: a slice of the mailing list does not get the offer, and the lift is the difference in purchase rate between the two. Customer experience measurement borrows the same idea and applies it to changes that are not campaigns: a new complaint-handling process, a rewritten onboarding sequence, a retrained team, a changed policy.
Three things are often mistaken for a control group and are not. A before-and-after comparison has no control, only a different time. A segment comparison (customers who complained versus those who did not) compares people who differ for reasons unrelated to the change. And a survey sample is a way of observing customers, not a way of assigning who gets what. A control group exists only when the people running the change decided, in advance, who would not get it and why those people are comparable.
The phrase “holdout group” means the same thing. It emphasizes the act of holding some customers back from a change that is otherwise going to everyone, which is how most CX holdouts work in practice.
Why before-and-after comparisons are a story, not evidence
Comparing this quarter to last quarter tells you what happened over time. It cannot tell you why, because everything else changed over the same time: the season, the economy, the marketing calendar, the product, the competitor. Any of them could explain the difference. Most of them will, if you look hard enough, and someone in the room always does.
A control group removes that problem by holding time constant. Two groups of customers live through the same quarter, the same season, the same competitor outage. One gets the new process. One does not. Whatever difference appears between them at the end is the part that the new process can claim.
The usual objection is that experience changes cannot be A/B tested. A new store layout, a retrained call center team, a rewritten policy: you cannot serve version B to half the people who walk through the door. That is true of the tools. The method does not care, and the role of measurement in keeping a program alive depends on using it anyway.
Five ways to hold out when the A/B tools cannot reach
You do not need a testing platform to have a control group. You need a way to split customers into two comparable sets and give one of them the change first.
A branch or a store. Roll the new process out in some locations and not others. Pick the locations so that the two sets look alike on the things that matter: size, customer mix, region, performance last year. Then compare.
A region. The same idea at a larger scale, useful when the change involves a policy or a team that is organized regionally anyway.
A cohort. Apply the change to customers who joined after a certain date and compare them with those who joined just before it. This works well for onboarding changes, and it lines up naturally with the leading indicators of lifetime value that show up in the first ninety days.
A staggered rollout. If the change is going everywhere eventually, roll it out in waves. Each wave that has not yet started is the control group for the waves that have. This is the least controversial option, because nobody is being denied anything, only scheduled.
A random ten percent. Where the change is something you control at the customer level (a proactive check-in, a new notification, a follow-up call after a complaint), assign a random slice to keep the old process for a defined period. Random is important. Choosing the ten percent by hand introduces the very bias you are trying to remove.
| Holdout unit | Best for | Main risk | Comparability comes from |
|---|---|---|---|
| Branch or store | Process and training changes delivered in person | Staff in holdout branches copy the new process informally | Matching on size, mix and last year’s results |
| Region | Policy changes and regionally organized teams | Regional events (weather, a local competitor) confound the result | Matching, plus a long enough period to average out events |
| Cohort by join date | Onboarding and early-relationship changes | Seasonality in who joins when | Adjacent cohorts, compared at the same relationship age |
| Staggered waves | Anything rolling out everywhere anyway | Later waves hear about the change and expect it | Random or matched ordering of waves |
| Random customer slice | Customer-level actions such as check-ins or follow-ups | Exclusion errors in sends and scripts | Randomization itself |
How to run a minimum viable control group
You can do this next quarter without a new tool.
- Write down the outcome and the timeframe before anything starts. “Retention at 90 days among customers who contacted support” is an outcome. “Better experience” is a hope. Use the same definitions your retention reporting uses, so the result can be compared with the numbers everyone already sees.
- Choose the unit of holdout: customer, branch, cohort or wave. Pick the smallest unit you can actually keep separate.
- Decide how big the holdout must be. Ask: is the difference I am hoping for bigger than the difference I would see between two random groups by chance? If nobody can answer that, ask someone who can before you start, or make the holdout larger than feels comfortable.
- Assign at random, or as close as you can get. For branches or regions, match on last year’s numbers and flip a coin for which side gets the change.
- Protect the boundary. Tell the people who need to know, exclude the control group from every related send and script, and check monthly that it is still clean.
- Compare only what you wrote down in step one. Resist the urge to hunt for some metric that moved. Something always moved.
A worked example (illustrative, round numbers)
A retailer with forty stores wants to know whether a new returns process reduces repeat complaints. It pairs the stores by size and last year’s complaint volume, giving twenty matched pairs, and flips a coin within each pair to decide which store gets the new process for one quarter. At the end of the quarter, the twenty treatment stores handled 4,000 returns and 200 of them generated a follow-up complaint, a rate of 5 percent. The twenty control stores handled 4,000 returns and 320 generated a follow-up complaint, a rate of 8 percent.
The three-point gap is the effect the new process can claim. Both groups lived through the same quarter, the same promotions and the same weather. Whether three points is enough to justify the cost of rolling the process out to all forty stores is a separate question, and it is now a question with a number in it.
What quietly breaks a control group
The method is simple. The execution has four traps, and I have fallen into each of them.
Contamination. The control group finds out. Staff in the holdout branch hear about the new process and start doing it informally. Customers in the control cohort get the new email because someone forgot to exclude them from a send. Once the two groups have been treated the same, the comparison is dead. Guard the boundary, and check it throughout, not just at the start.
Groups too small to see anything. If the effect you are hoping for is a couple of points of retention and your holdout is 80 customers, you will not see it, and neither will anyone else. The size you need depends on the size of the effect you expect and how much the outcome naturally varies. A holdout that is too small does not just fail to prove the change worked; it produces a random result that someone will present as if it meant something.
Comparing volunteers with non-volunteers. Customers who opted into the pilot are different from those who did not, in ways that have nothing to do with the pilot. So are branches whose managers put their hands up. Enthusiastic people produce better numbers regardless of the process they are running. If the treatment group chose itself, you have measured enthusiasm. This is a recurring problem in pilot programs, and building buy-in for a pilot has to be done without letting the enthusiasts pick the sites.
Hunting for the metric that moved. If the outcome you wrote down did not move, the temptation is to look at twelve others until one did. With twelve metrics, one will have moved by chance. Report the pre-registered outcome first, and label anything else as exploratory.
When a control group is not the answer
There are cases where holding out is wrong, and cases where it is impossible, and it is worth knowing which is which.
Never hold out safety, refunds or the fixing of a genuine error. Those go to everyone, immediately, and you find another way to measure. The ethical question about holdouts is real, and the answer is that a brief, bounded holdout of an improvement is a small cost that buys the evidence to roll the improvement out properly and keep it funded once the initial excitement fades, which is the fate of most pilots that cannot prove themselves (Corporate attention deficit and why pilot programs fail is about that fate). But an improvement is not the same as a correction, and corrections are not withheld.
When the change cannot be split. A rebrand, a new pricing structure, a company-wide policy that customers would compare across the boundary. Here the fallback is a before-and-after with the confounders written down in advance: list the other things that are changing over the same period, and say how you will tell their effect apart, before the results arrive.
When the customer base is too small. A business with a few hundred customers cannot split them and expect to see a modest effect. Use customer-level leading indicators instead, and be honest that the evidence is suggestive.
When the outcome takes years. A holdout on something whose effect appears in year three is a holdout nobody will maintain. Measure the early indicator with the control group, and let the lagging outcome confirm it later.
A control group will not make a weak change look strong. That is the point. What it does is let a strong change look strong to the people who decide whether it continues.
Where to start
- Pick the next change on the CX roadmap that is going to more than one place. A process, a script, a check-in, a policy. Rollouts that are already planned in phases are the easiest.
- Write the outcome, the definition and the timeframe on one page. Get the person who owns the customer analysis to agree the definition before anything starts.
- Choose which wave, branch or slice goes last, and make it random or matched. Treat that group as the most important one in the plan.
- List every send, script and system that would need to exclude the holdout. Do the exclusions, then check them at the end of the first month.
- Put a date in the calendar for the comparison. One number, the one you wrote down, for each group, on one slide.
FAQ
What is a control group in customer experience measurement?
A control group is a set of customers, branches or cohorts that deliberately does not receive a change, chosen so that it is comparable to the group that does. Because both groups live through the same period, the difference in outcome between them is the effect of the change rather than of the season, the economy or a competitor. It is the only way to say a customer experience change caused a result rather than merely preceded it.
What is the difference between a control group and a holdout group?
They are the same idea with different emphasis. “Holdout group” stresses that some customers are being held back from a change that is otherwise going to everyone, which is how most customer experience holdouts are set up. “Control group” is the general term from experimental design and covers any comparable group that does not receive the treatment.
How big does a control group need to be?
Large enough that the effect you expect is bigger than the difference two random groups would show by chance. That depends on the size of the expected effect and how much the outcome naturally varies, so there is no single number. If nobody on the team can estimate it, ask someone with statistical training before starting, or make the holdout larger than feels comfortable.
Is it ethical to withhold an improvement from a control group?
A brief, bounded holdout of an improvement is generally defensible, because it produces the evidence needed to roll the improvement out properly and keep it funded. Safety fixes, refunds and corrections of genuine errors should never be withheld from anyone. Where possible, stagger the rollout so that nobody is denied the change, only scheduled to receive it later.
Can you use a control group without an A/B testing tool?
Yes. A control group is a method, not a piece of software. Rolling a change out to some branches and not others, to customers who joined after a date and not before, in staggered waves, or to a random slice of customers held on the old process all create a control group without any testing platform.
How long should a holdout last?
Long enough for the outcome you wrote down to be observable, and no longer. For an onboarding change measured by ninety-day retention, that is about a quarter. Holdouts that run for years are rarely maintained and usually become contaminated, so measure an early indicator with the holdout and let the lagging outcome confirm it later.