How the signal stays honest
A satisfaction score is the number everybody reports and the least useful thing this system produces. Two other numbers determine whether it means anything at all, and a third determines whether the follow-up is worth the effort.
Key takeaways
- Group into a fixed theme list, for the same reason as search queries: comparability.
- Response rate governs everything. A score from 4% of customers describes those 4%.
- Measure recovery: did the people you followed up with buy again?
- Report the theme counts, not just the score. Themes are what somebody can fix.
- Resist tying the score to anybody’s pay, or the ask itself will be gamed.
Fixed themes
The same argument as the search rank reporter: a model asked to cluster comments will produce sensible groups that differ slightly between runs, and every trend built on them is meaningless. A fixed list — delivery, product quality, communication, price, staff, website — produces the same groups forever.
Six or seven themes is right. The model’s job is classification into that list or into unmatched, and a growing unmatched bucket is the signal that the list needs a new entry, reviewed quarterly.
Response rate governs everything
- App integration
- Machine learning
- Security & identity
- Management
- Analytics
Who does not answer
The people who ignore a satisfaction survey are not a random sample. They skew towards the mildly satisfied — nothing went wrong, nothing was memorable — and away from both extremes. A score built from the people who did answer is a score of the people who felt something.
That does not make it useless; it makes it a different measurement from the one people assume. Reporting the rate next to the score is the whole mitigation, and it costs one extra number.
Does following up work
- Management
- Analytics
- Front-end & mobile
- People
The comparison is imperfect — two is not a control group — and it is the right question, which is more than most feedback reporting manages. Seventeen of twenty-nine unhappy customers ordering again is a recovery rate, and it is the number that justifies somebody’s ninety seconds.
Over a year the sample gets large enough to mean something, and the trend within it is more informative than the level: a recovery rate that falls after somebody changed how replies are written is a real finding.
The pressure to game it
What happens when the score becomes a target
- The ask gets selective. Surveys stop going to orders that went badly, which is the easiest and most invisible manipulation available.
- The wording gets leading. “We hope you were happy with your order — how did we do?” scores measurably higher than a neutral ask.
- The timing moves. Asking immediately after a pleasant interaction rather than after the outcome is known.
- The counter-measure is to report the ask rate alongside the score: how many eligible orders were surveyed. A score that rose while the ask rate fell is not a score that rose.
This is the strongest argument against tying the score to anybody’s pay or targets, and it is worth making explicitly to whoever suggests it. The measurement is only useful while nobody has a reason to move it, and it is very easy to move in ways that leave no trace in the score itself.
Reporting the ask rate is the cheapest available defence: it makes the most common manipulation visible in the same table as the number it was meant to improve.
Next: what all of this costs to run.
All posts