Customer Experience

Why pilot programs fail: corporate attention deficit and the fix

Why pilot programs fail: rarely from bad results, usually because attention runs out before anyone decides. Five structures that make a pilot survive.

Table of contents
  1. Key takeaways
  2. Why pilot programs fail: attention, not evidence, is the scarce resource
  3. The five ways a pilot dies
  4. The pilot graveyard and the pilot-to-scale gap
  5. How to design a pilot that survives its own organization
  6. Pilot vs phased rollout vs experiment: which to use when
  7. What quietly makes pilot programs fail even with the structure in place
  8. When a pilot is the wrong instrument
  9. Where to start
  10. FAQ

Most companies I have worked with have a folder somewhere called “Pilots”. Inside are subfolders with hopeful names and dates two or three years old. Each contains a kickoff deck, a plan, a few weeks of data, and then nothing. No final report. No decision. Just an absence, as if everyone stepped out for lunch and never came back. If you want to know why pilot programs fail, that folder is the honest answer.

Ask what happened to any one of them and you rarely hear “it did not work”. You hear “the sponsor moved to another division”, or “that was before the reorg”, or “we got pulled onto the migration”, or, most honestly, “I am not sure, actually”. The pilot did not fail. It was forgotten.

A pilot program is a limited, time-boxed trial of a change, run on a subset of customers or teams so that the company can decide whether to scale it. The decision is the point. A pilot that ends without one has not been run; it has been started.

Key takeaways

  • Pilot programs rarely die from bad results; they die because organizational attention runs out before anyone is obliged to decide anything.
  • Every unfinished pilot leaves a residue in the company, and “we tried that” becomes an excuse even though nothing was tried to a conclusion.
  • The pilot-to-scale gap is the cost of turning every exception the pilot ran on into a rule with an owner, and it has to be costed before launch.
  • Five structures, agreed before the first customer is touched, let a pilot survive without anyone remembering to care: a decision rule, a named owner, a date, a comparison group and a costed scaling answer.
  • A pilot is the wrong instrument when the decision is already made, when the change is cheap to reverse, or when the population is too small to compare.

Why pilot programs fail: attention, not evidence, is the scarce resource

I call this corporate attention deficit, and I should say plainly that the phrase is borrowed from a real condition that real people live with; it is a metaphor about organizations, not a joke at anyone’s expense. Companies, unlike people, can choose to fix theirs, and mostly do not.

The pattern is consistent. A pilot begins with a burst of attention: an executive interested, a team energized, a kickoff meeting with snacks. Then the quarter changes. The executive’s priorities change with it, because that is what quarters are for. The team’s manager needs them back. The dashboard that IT was going to build is still a ticket. By month three the pilot has a caretaker rather than an owner, and by month five the caretaker has a new job.

Notice what is missing from that story: results. The pilot did not produce a bad number. It never produced a number at all, or it produced one that nobody had agreed in advance would mean anything. There was no moment when someone was obliged to say “scale it” or “stop it”, so nobody said either, and a thing that nobody stops or scales simply thins out.

Most of the remedy is the internal selling work described in the post on building buy-in for a pilot. The difference is the audience. Buy-in is sold to the company as it is today. Survival has to be sold to the company as it will be three quarters from now, with a different set of worries.

The five ways a pilot dies

The stories in the “Pilots” folder sound different and follow the same few scripts.

How it dies What it sounds like afterwards When it usually happens The structure that was missing
The sponsor moves on “That was her project” Month two to four A named owner whose handover starts with the pilot
The team is reclaimed “We got pulled onto the migration” Month two Time agreed with the manager in writing, not borrowed
The number means nothing “Retention was up, but the market was up too” At the first review A comparison group held out from the start
No decision moment “We were going to review it in Q3” Never, which is the point A decision date already in twelve calendars
A result with no bridge “It worked, but we could not resource the rollout” After a good result A costed answer to “what would scaling take”

Only the last row involves a result at all. The other four are pilots that ended without ever testing the idea they were built to test.

The pilot graveyard and the pilot-to-scale gap

Two things follow from this, and both cost more than the pilots themselves.

The first is the graveyard. Every unfinished pilot leaves a residue: a team that learned its effort was disposable, a frontline that was asked to change how it worked and then never told it could stop, and a set of customers who were called once, with care, and then never again. The next pilot inherits all of it. “We tried that” becomes an excuse, even though nothing was tried to a conclusion.

The second is the pilot-to-scale gap. Even the pilots that survive to a result face a chasm between “it worked for a couple of hundred customers with a volunteer team” and “it runs for everyone as part of the job”. A pilot is built out of exceptions: borrowed people, a spreadsheet instead of a system, a manager quietly waiving the call-time target. Scaling means turning each exception into a rule, and each of those rules has an owner who was not at the kickoff. If nobody costed that crossing before the pilot started, the result arrives with no bridge to carry it over, and it joins the graveyard anyway, with a better epitaph.

How to design a pilot that survives its own organization

The remedy is not more enthusiasm. Enthusiasm is what the attention deficit feeds on. The remedy is structure, agreed before launch, that does not depend on anyone remembering to care. Five pieces, in the order to write them.

  1. A pre-agreed decision rule. In writing, before the first customer is touched: if the treated group’s repeat rate after twelve weeks is higher than the comparison group’s by at least the agreed margin, we scale; if it is lower, we stop; if it is in between, we run one more quarter and then decide. The numbers can be argued about. The existence of the rule cannot, because without it the pilot’s result is an opinion.
  2. A named owner. One person, whose name is on the plan and in the sponsor’s calendar. Not a team, not a working group. Someone who will be asked “how is your pilot going” and cannot answer “you would have to ask the group”. If that person changes jobs, the first item in their handover is the pilot.
  3. A decision date on the calendar. An actual meeting, booked now, with the sponsor and the people who would have to fund scaling, at which the decision rule is applied. Not “we will review in Q3”. A date, a room, a calendar invite that survives the reorg because it is already in twelve diaries.
  4. A comparison group. Held out from the start, chosen the same way as the treated group, so that on the decision date the conversation is “here is the difference” rather than “well, the market was also up”. This is the part of measurement people most often skip and most often regret, and it is what makes the decision rule enforceable. The wider measurement plan, the three questions and the one-page monthly report, is covered in its own post.
  5. A costed answer to “what would scaling this take”. Before launch, not after. Which exceptions would become rules. Who owns each one. What it would cost in people, systems and management time. The number will be rough, and it should be written down anyway, because the alternative is discovering on the decision date that nobody knows whether “yes” is affordable, and postponing, which is how a pilot with a good result still dies.

A worked example (illustrative)

The figures are round and invented, to show what a written decision rule looks like.

A callback pilot treats 1,800 customers whose tickets closed without a resolution and holds out 200 at random. The rule, agreed in writing before launch: at twelve weeks, if the treated group’s repeat-purchase rate is at least five points higher than the held-out group’s, the head of service will fund two permanent roles and the pilot becomes a standing process; if it is lower than the held-out group’s, it stops that week; anything in between earns one more quarter, after which the same rule applies with no further extension.

The scaling cost is written on the same page: two roles, a change to the ticket-closure workflow so the list generates itself, and a monthly hour of the analyst’s time. On the decision date, 30 in 100 treated customers have bought again against 24 in 100 held out. Six points. The rule says scale, the cost is already known, and the meeting takes ten minutes. Nobody had to remember to care. They only had to turn up to a meeting that was already in their calendar.

Pilot vs phased rollout vs experiment: which to use when

“Pilot” gets used for three different things, and the confusion is part of why pilots fail. Each has a different purpose and a different ending.

Pilot Phased rollout Controlled experiment
Purpose Decide whether to scale Scale something already decided, safely Learn which of two options works better
Comparison A held-out group The regions not yet reached Random assignment between options
Size Small enough to run on exceptions Everyone, eventually As large as the difference needs
Ends with Scale, stop or change Full coverage A finding
Fails when Nobody decides A phase reveals a problem nobody planned for The groups were not really random
Use it when The idea is unproven and the cost of being wrong is real The decision is made and the risk is in execution Two plausible designs and the means to assign at random

If the decision has already been made, run a phased rollout and call it that. Dressing a rollout up as a pilot invites a review that nobody intends to act on. If two versions of the idea are competing, run an experiment. A pilot is for the specific case where the company genuinely does not know whether to do this at all.

What quietly makes pilot programs fail even with the structure in place

The decision date moves, once. The sponsor is traveling; the meeting slides a month. A date that has moved once can move again, and a pilot with a movable date has no date. Hold the meeting without the sponsor if necessary, and apply the rule.

The owner becomes a working group. It starts with one name and ends with five, because the pilot touches five teams. Five owners are none. Keep one, and let the others be consulted.

The comparison group leaks. An agent sees a held-out customer with an obvious problem and calls them anyway. After a quarter of kindness, the two groups are one group, and the decision rule has nothing to compare. Protect the list and explain why.

Sales hears about it from a customer. In any company with account teams, a pilot that contacts customers is touching someone’s accounts. Whether sales cares about the program is settled in the week before launch, and a pilot that skipped that week has an objector it did not need.

When a pilot is the wrong instrument

When the change is cheap and reversible. Rewording an automated email or removing one redundant field from the renewal form does not need a quarter, a hold-out and a steering group. Do it, watch the number for a month, and undo it if it hurt.

When the population is too small to compare. A pilot that touches forty customers a quarter will not produce a difference anyone can trust. Run it as a service standard instead, measure whether the action happened, and collect the customers’ own words.

When the organization is mid-upheaval. A merger, an acquisition or a reorganization changes sponsors, budgets and attention at once, and the people who might have owned the pilot are wondering about their own roles. Shrink the pilot to something one team can run with no money, or wait a quarter. Selling harder into a distracted company only adds to the graveyard.

Where to start

  1. Open the “Pilots” folder and write one line per subfolder: what it tested, whether a decision was ever made, and what it cost. That page is the case for doing the next one differently.
  2. Write the decision rule for the next pilot in three sentences: scale if, stop if, extend once if.
  3. Name one owner and put the pilot at the top of their handover template today, before they need it.
  4. Book the decision meeting with the sponsor and the budget holders, on a date that is already in their calendars.
  5. Cost the scaling answer on half a page: which exceptions become rules, who owns each, and what it would take in people and systems.

FAQ

Why do pilot programs fail?

Most pilot programs fail because organizational attention runs out before anyone is obliged to decide anything, not because the results were bad. The sponsor moves, the team is reclaimed, the number was never agreed to mean anything, or no decision meeting was ever booked. The pilot thins out rather than ending, and “we tried that” becomes an excuse even though nothing was tested to a conclusion.

How long should a pilot program run?

Long enough for the agreed outcome to show a difference between the treated group and the comparison group, and no longer. For a customer callback or retention pilot that is usually one quarter, with a leading indicator checked monthly and a single pre-agreed extension of one more quarter if the result is borderline. A pilot with no end date is not a pilot.

What is the pilot-to-scale gap?

The pilot-to-scale gap is the distance between a pilot that worked for a few hundred customers with a volunteer team and the same idea running for everyone as part of the job. A pilot runs on exceptions such as borrowed people, spreadsheets and waived targets, and scaling means turning each exception into a rule with an owner. Costing that crossing before launch is what stops a successful pilot from being shelved anyway.

How do you decide whether to scale a pilot?

Apply a decision rule that was written before launch: scale if the treated group beats the comparison group by at least the agreed margin, stop if it does worse, and extend once if the result is in between. Hold the decision at a meeting booked in advance with the sponsor and the people who would fund scaling. Have the cost of scaling already on the page, so the decision is about the result and not about whether “yes” is affordable.

What is corporate attention deficit?

Corporate attention deficit is a metaphor for the way organizations lose interest in an initiative as quarters change, sponsors move and priorities shift, so that pilots are abandoned rather than concluded. The phrase borrows from a real condition that real people live with and is not a comment on anyone. Unlike people, companies can fix theirs with structure: a decision rule, a named owner, a date, a comparison group and a costed scaling plan.

Related articles