Skip to content

Part 5 of 7 · Feedback follow-up router series ~5 min read

How the signal stays honest

A satisfaction score is the number everybody reports and the least useful thing this system produces. Two other numbers determine whether it means anything at all, and a third determines whether the follow-up is worth the effort.

Key takeaways

  • Group into a fixed theme list, for the same reason as search queries: comparability.
  • Response rate governs everything. A score from 4% of customers describes those 4%.
  • Measure recovery: did the people you followed up with buy again?
  • Report the theme counts, not just the score. Themes are what somebody can fix.
  • Resist tying the score to anybody’s pay, or the ask itself will be gamed.

Fixed themes

The same argument as the search rank reporter: a model asked to cluster comments will produce sensible groups that differ slightly between runs, and every trend built on them is meaningless. A fixed list — delivery, product quality, communication, price, staff, website — produces the same groups forever.

Six or seven themes is right. The model’s job is classification into that list or into unmatched, and a growing unmatched bucket is the signal that the list needs a new entry, reviewed quarterly.

Response rate governs everything

How a satisfaction score is qualified before being reportedA vertical chain of five steps entered by a box labelled A reported score, say eight point one. Step one asks what response rate it came from; under ten per cent exits to a note that it describes only that ten per cent. Step two asks who did not answer and whether they are systematically different. Step three asks whether anonymous responses are mixed in, since they skew low; if so it exits to Report separately as two numbers. Step four asks whether there were enough responses this period against a floor; if not it exits to Do not report a change and state the volume instead. Step five is a score worth quoting, with its rate beside it. A note says the score is never reported without the response rate next to it, ever.AWS ACCOUNTA reported scoresay 8.1From what response rate?Under 10%it describes the 10%lowWho did not answer?systematically different?Anonymous mixed in?they skew lowReport separatelytwo numbersyesEnough responses this period?a floorDo not report a changesay the volumenoA score worth quotingwith its rate beside itThe score is never reported without the response rate next to it. Ever.
Fig 1. Why a score is never quoted alone. Each of these checks is a way the headline number can be true and misleading at the same time.
  • App integration
  • Machine learning
  • Security & identity
  • Management
  • Analytics

Who does not answer

The people who ignore a satisfaction survey are not a random sample. They skew towards the mildly satisfied — nothing went wrong, nothing was memorable — and away from both extremes. A score built from the people who did answer is a score of the people who felt something.

That does not make it useless; it makes it a different measurement from the one people assume. Reporting the rate next to the score is the whole mitigation, and it costs one extra number.

Does following up work

Whether following up on unhappy responses changes behaviourA horizontal row of five boxes. Unhappy responses: thirty-one this quarter. Followed up: twenty-nine within the hour. Ordered again: seventeen. Unhappy and not followed up: two, both by accident. Ordered again: none. A note says the numbers are small and it is still the number to watch, because recovery matters rather than response count.THE ONLY MEASUREMENT THAT JUSTIFIES ANY OF THISUnhappy responses31 this quarterFollowed up29 within the hourOrdered again17Unhappy, not followed up2, both by accidentOrdered again0Small numbers, and it is still the number to watch. Recovery, not response count.
Fig 2. The measurement that decides whether the follow-up is worth doing. It is a small sample and it is the right question, which most feedback reporting never asks.
  • Management
  • Analytics
  • Front-end & mobile
  • People

The comparison is imperfect — two is not a control group — and it is the right question, which is more than most feedback reporting manages. Seventeen of twenty-nine unhappy customers ordering again is a recovery rate, and it is the number that justifies somebody’s ninety seconds.

Over a year the sample gets large enough to mean something, and the trend within it is more informative than the level: a recovery rate that falls after somebody changed how replies are written is a real finding.

The pressure to game it

What happens when the score becomes a target

  • The ask gets selective. Surveys stop going to orders that went badly, which is the easiest and most invisible manipulation available.
  • The wording gets leading. “We hope you were happy with your order — how did we do?” scores measurably higher than a neutral ask.
  • The timing moves. Asking immediately after a pleasant interaction rather than after the outcome is known.
  • The counter-measure is to report the ask rate alongside the score: how many eligible orders were surveyed. A score that rose while the ask rate fell is not a score that rose.

This is the strongest argument against tying the score to anybody’s pay or targets, and it is worth making explicitly to whoever suggests it. The measurement is only useful while nobody has a reason to move it, and it is very easy to move in ways that leave no trace in the score itself.

Reporting the ask rate is the cheapest available defence: it makes the most common manipulation visible in the same table as the number it was meant to improve.

Next: what all of this costs to run.

All posts