Organisations that move survey operations to a voicebot tend to see the same pattern in the first quarter. Contact volume rises noticeably and the satisfaction score falls. The usual interpretation is that customers dislike the automated system.
Usually something else is happening. Customers who never took part before are now taking part, and that group was already less satisfied. The score did not fall. The measurement became more representative.
A measurement system that cannot tell these two apart will send improvement work in the wrong direction. So the methodology conversation has to come before the questionnaire.
Non-respondents carry as much information as respondents
Manually dialled surveys reach a small slice of the target population, and only some of those complete the call. Those two filters quietly shape the sample. Customers who can take a call during working hours, who have time to spare, and who already have a good relationship with the brand are over-represented.
Automated calling loosens both filters. Evening slots, shorter interviews and repeat attempts reach a broader group, and the score typically drops a little as a result. That drop is not a problem; it is the bias in the previous measurement becoming visible.
The right response is to run both methods in parallel during the transition. Measure the same segments with the old and the new method for at least one quarter, quantify the difference, and you can read the trend correctly afterwards. Organisations that skip this step spend the following year unable to make any year-on-year comparison.
What exactly do you mean by response rate
"Our response rate is 60 per cent" comes up in almost every review, and it means something different nearly every time. Whether it is calculated on numbers dialled, people reached, or interviews started is usually left unstated.
AAPOR's widely adopted standard definitions separate these: response rate, contact rate, cooperation rate and completion rate are different quantities (AAPOR Standard Definitions). Writing down which definition your organisation uses ends an argument you would otherwise have three months later.
What NPS actually tells you
Net Promoter Score entered mainstream use through Reichheld's 2003 Harvard Business Review article, with a strong claim attached: the "would you recommend" question was the single best predictor of growth (Reichheld, HBR 2003).
That claim has been seriously contested since. Keiningham and colleagues, publishing in the Journal of Marketing in 2007, found NPS no better at predicting firm revenue growth than other satisfaction and loyalty metrics. Later replication work noted that the original finding rested on cross-sectional data and did not hold up as strongly under longitudinal analysis (replication summary, MeasuringU).
None of this means you should drop NPS. It remains a workable way to build a shared language around a single number. But "if we lift NPS by five points, revenue grows by X" is not a claim the literature supports. Use NPS as a management instrument, not as a forecasting model.
For the same reason, collecting NPS alone is not enough. Without the answer to "why", there is nothing actionable in your hands.
Effort may be a better operational signal than satisfaction
Dixon and colleagues argued in 2010 that trying to delight customers has a weaker effect on loyalty than commonly assumed, and that reducing the effort customers expend is more decisive (Dixon, Freeman, Toman, HBR 2010).
Operationally this matters because the effort question, how hard did you have to work to get this resolved, points at something more concrete than satisfaction does. A low CSAT reports a feeling. A high effort score usually reports a broken process.
The effort question is also easy to ask by voice: short and one-dimensional.
Design constraints specific to the voice channel
Question sets copied from web surveys do not survive a phone call. A few concrete differences.
Scale length. Reading a 0 - 10 scale aloud takes time and the respondent forgets the anchors. A 1 - 5 scale performs noticeably better by voice. But changing scale breaks comparability with your historical data, which makes it one of the more expensive decisions in a measurement system. If it has to change, run both scales in parallel through the transition.
Question order. Earlier questions influence later ones. Ask the overall satisfaction question after a series of specific problems and the score comes down. Fix the order once and do not vary it between periods.
Where the open-ended question goes. Put it first and completion drops. Put it last and more people reach it, and they explain a score they have already given, which makes coding easier.
Interview length. Drop-off climbs quickly past three minutes on the phone. Eight questions is usually the practical ceiling. If you need more, splitting them across periods works better than asking them all at once.
Changing channel changes the score by itself
This is well established in survey methodology: the same question asked in a different mode gets a different answer. People speaking to a human interviewer tend to give softer scores than they give an automated system, while automated channels can elicit more candid answers on sensitive topics. These effects are covered in the literature under mode effects (see Dillman, Smyth and Christian, Internet, Phone, Mail, and Mixed-Mode Surveys).
The practical consequence: in the quarter you change channel, do not compare scores. Calibrate first.
Consistency in coding open-ended answers
The most valuable output of a voice survey is not the score, it is the open-ended responses. Labelling those automatically is feasible, but it raises a quality question of its own: does the same answer get the same label at two different times?
The practical way to check is to human-code a small sample regularly and compare against the automatic labels. When agreement falls, the usual cause is a label set that has grown too large. Consistency degrades quickly in systems with fifty topic labels. Around fifteen main categories with free keywords underneath is a more sustainable structure.
Transcript quality feeds label quality directly. If product names, branch names and campaign codes are not recognised, open-ended analysis fails systematically in the same places. Adding those terms to a lexicon delivers far more than swapping models.
Setting up the measurement system
- Which contact types trigger measurement, and within how many hours
- How is the sample drawn, are there segment quotas, will results be weighted
- Which definition of response rate gets reported
- Are scale and question order fixed, and is there a parallel measurement window if they change
- How many categories are in the open-ended label set, and who owns it
- How often are automatic labels compared against human coding
- When score changes are reported, is the change in sample composition reported alongside
The last point is the one most often skipped. A report that shows a score without the period's sample composition invites the wrong improvement decision.
To see how voice survey flows are built and how results are broken down, look at Survey and Customer Experience Measurement.
References
- Reichheld, F. (2003). The One Number You Need to Grow. Harvard Business Review. hbr.org
- Keiningham, T. L., Cooil, B., Andreassen, T. W., Aksoy, L. (2007). A Longitudinal Examination of Net Promoter and Firm Revenue Growth. Journal of Marketing, 71(3). Summary and replication discussion: measuringu.com
- Dixon, M., Freeman, K., Toman, N. (2010). Stop Trying to Delight Your Customers. Harvard Business Review. hbr.org
- AAPOR, Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys. aapor.org
