Exit survey analysis: turning free text into decisions
Exit-survey data becomes decision-grade later than most merchants think. At 100 coded responses, a reason sitting at 30% carries a margin of error of about nine points either way; at 25 responses, eighteen points. So: code free text monthly against a fixed codebook, act on offer routing at about 100 responses, hold pricing and product decisions for about 200, and treat every stated reason as a claim to verify against behavior, not a fact.
The collection side is settled - one required question, optional open text, per the one-question rule. This post is about the other end of the pipe: what to do with the answers once you have them, and when you have enough of them to do anything at all.
What the survey layer is actually for
An exit survey has one job in real time - routing the save offer - and one job in aggregate: telling you which problem to fix next. The routing job needs no analysis at all; the reason bucket fires a rule and the flow responds, per the taxonomy and routing map. The aggregate job is where analysis lives, and it is where most stores either do nothing or do too much too early.
"Nothing" looks like a wall of free text nobody reads. "Too much too early" looks like a pricing change justified by eleven angry responses. Both failure modes have the same cure: a boring, repeatable coding routine and a hard rule about sample sizes.
Coding free text without a data team
Coding means assigning each free-text response to a category from a fixed list, and a solo merchant can do a month's worth in half an hour. The method that works at WooCommerce scale:
- Start the codebook from your reason buckets. The radio options are your top-level codes. Free text mostly elaborates the selected bucket, so you are usually assigning subcodes: "too expensive" splits into "budget shrank", "price vs value", "found cheaper".
- Code the primary complaint only. A response that mentions price and a shipping problem gets one code - the one the customer leads with. Multi-coding feels rigorous but makes every percentage ambiguous.
- One coder, one pass, monthly. Consistency beats sophistication. The same person applying slightly wrong rules the same way every month produces usable trend data; three people applying good rules differently produce noise.
- Quote-tag as you go. Paste the two or three most vivid verbatims per code into the monthly note. These do more to move product decisions than the percentages do.
- Add a new code only under pressure. A code earns its place when five or more responses in a quarter clearly demand it. Below that, leave them in "other".
The tooling is a spreadsheet, on purpose. Six columns cover it: cancel date, stated reason, subcode, verbatim worth quoting (yes/no), tenure at cancel, and last-activity date. The last two columns look like overkill for a survey log; they are what makes the stated-vs-revealed checks below possible without a second export.
Even the biggest taxonomy in the public data leaks: Churnkey's State of Retention 2025 shows 17.85% of 3 million cancellation sessions landing in "Other reasons". Your codebook will leak too. The goal is a leak you monitor, not a leak you ignore.
Minimum sample sizes before acting
The margin of error on a reason share is plain binomial arithmetic, and it is worse than intuition suggests. For a reason that truly sits at 30% of cancels, the 95% confidence interval by number of coded responses:
| Coded responses | Margin of error | You can distinguish |
|---|---|---|
| 25 | about +/- 18 points | almost nothing |
| 50 | about +/- 13 points | dominant vs minor reasons |
| 100 | about +/- 9 points | the top two or three ranks |
| 200 | about +/- 6 points | real mix shifts |
| 400 | about +/- 4.5 points | most changes worth acting on |
These are not benchmarks from a vendor report; they fall out of the standard formula, and they assume clean, involuntary-churn-excluded data. The practical thresholds we draw from them:
- Reorder or re-target offers at ~100 coded responses. Getting the top bucket right only requires rank order, and ranks stabilize around there.
- Make pricing or product decisions at ~200. A 32% price share and a 24% price share tell very different stories, and below 200 responses you cannot tell them apart.
- Call a trend only with ~350 per period. Detecting a genuine ten-point shift between two months (say 30% -> 40%) with conventional statistical power needs roughly 350 responses in each period. A store with 40 cancels a month should compare quarters or halves, never months.
For a store doing 30-50 cancels a month, this timeline is sobering: rank-order confidence arrives in a quarter, decision confidence in half a year. That is the honest cost of small scale, and pretending otherwise just launders noise into strategy.
One refinement: the thresholds assume decisions of medium weight, and the bar should scale with the cost of being wrong. Reordering which offer shows first is cheap and reversible - act at 100, or even 75, because a mistake costs you one quarter of slightly worse routing. Rewriting flow copy is nearly free - act whenever the verbatims are unanimous. Restructuring pricing or killing a product line is expensive and sticky - that deserves 200-plus responses and a second period saying the same thing. The question is never "is the data perfect" but "is the evidence stronger than the cost of acting wrongly".
Stated vs revealed reasons
A stated reason is what the customer picks; a revealed reason is what their behavior says; when the two disagree, believe the behavior. Churnkey's report makes the case from the stated side: analyzing freeform follow-ups at scale, they found "budget limitations" frequently serving as a repository for product frustration, disillusionment, and bad experiences - price is simply the easiest thing to type.
Three cross-checks reveal the real reason, and all three use data WooCommerce already has:
- Usage before cancel. Pull last order date, last login, or last download for each cancel. A "too expensive" from an active customer is a price cancel; a "too expensive" from someone dormant for two months is a usage or fit cancel that will not respond to a discount.
- Which offer they accepted. A customer who says "too busy" and then takes a discount over a pause revealed a price sensitivity the survey missed. In aggregate, acceptance mixes are informative but confounded: Churnkey's 2025 report has discounts at 53.9% of accepted offers versus 19.2% for pauses, which partly reflects what customers want and partly reflects what merchants most often present. Your own accept-given-shown rates are the clean version of this signal.
- Winback response. Someone who claimed "no longer need it" but returns on a plain 10% winback offer needed it fine; the price was the issue. Tag conversions from your winback sequence by original stated reason and the lies surface on their own.
The point is not that customers are dishonest - it is that a radio button is a low-bandwidth channel under social pressure. Stated reasons are a hypothesis generator. Behavior is the test.
In-flow answers and email answers are different datasets
If you also run a post-cancel follow-up email, keep its responses in a separate column and never pool the two streams into one percentage. They measure different populations. The in-flow required question is answered by essentially everyone who completes a cancellation, which is what makes its shares usable as prevalence estimates. The 48-hour email is answered by a self-selected sliver - disproportionately the customers with strong feelings in either direction, and disproportionately the articulate ones.
That does not make the email stream worthless; it makes it a different instrument. Use the in-flow stream for every number you compute - shares, trends, thresholds. Use the email stream for depth: it produces the paragraph-length explanations the in-flow textarea rarely gets. When the two disagree on prevalence ("half the emails mention shipping, but shipping is 6% in-flow"), believe the in-flow number and treat the email cluster as a lead worth investigating, not a share.
The "Other" share is a health metric
Track the percentage of responses coded "other" every month; it is the single best indicator of whether your taxonomy still fits your store. Churnkey's aggregate sits at 17.85%, and that is with a seven-category list tuned over millions of sessions. Reasonable working rules:
- Under 15%: healthy. Mine the free text quarterly for candidate codes anyway.
- 15-25%: your list is missing a bucket your customers need. The fix is usually adding a life-event or situation-changed option, per the taxonomy post.
- Over 25%: the options are worded in your language instead of the customer's. Rewrite them from verbatims, not from your org chart.
A rising "other" share with stable volume is an early warning that something new is happening - a competitor launch, a shipping change, a broken integration - before any named bucket moves.
From codes to decisions
Every coded batch should end by checking four signals against their evidence bars. The matrix we use:
| Signal | Minimum evidence | Action |
|---|---|---|
| One bucket over 40% of coded cancels | ~100 responses | Re-weight offers and flow copy toward that bucket |
| Price-stated cancels with dormant usage | ~30 matched cases | Fix onboarding or usage nudges, not pricing |
| "Other" share above 20% | 2 consecutive months | Rework the reason list wording |
| New subcode appearing repeatedly | 5+ in a quarter | Promote to codebook; alert product/ops |
Note what is absent: no signal in an exit survey justifies a price increase or decrease on its own. Survey data tells you where to look; the pricing decision needs revenue math from your actual cohorts, which is a different post and a different dataset.
The monthly 30-minute routine
Everything above compresses into a routine you can hold on the first Monday of each month:
- Export last month's cancels with reason, free text, tenure, and last-activity date. The ChurnStop dashboard exports this as one CSV; any flow tool that cannot is hiding your own data from you.
- Exclude involuntary churn from the export entirely - failed payments never answered the survey and belong in a different lane.
- Code the free text against the codebook. Cap it at 30 minutes; consistency beats depth.
- Update three numbers: bucket shares, "other" share, and running total of coded responses since the last decision.
- Check the decision matrix. If no threshold is met, write one sentence ("n=143, price still leads, no action") and stop.
The discipline of writing "no action" is the whole game. Exit-survey analysis fails in two directions, and the routine guards both: it guarantees the reading happens, and it makes acting early a visible violation of your own rule.
