Your QA sample rates agents well.
It finds themes badly.
Both jobs use the same recordings, so it is reasonable to assume one sample serves both. It does not, and the reason is arithmetic rather than opinion.
9 September 20266 min read
Two jobs that look like one
A contact centre listens to calls for two different reasons, and almost every quality programme is built for the first one only.
The first job is rating people. Did the agent greet correctly, disclose the price, agree a next step? The unit of analysis is the call and the subject is the agent. For that, a sample is not a compromise — it is the right tool. If you want a stable estimate of how an agent performs, a modest number of their calls gives it to you, and listening to more of them adds cost without adding certainty.
The second job is finding out what customers are saying. The unit is the whole corpus and the subject is the customer. Here a sample is not a smaller version of the answer. It is a different answer.
The arithmetic, with round numbers
Take an operation handling 39,000 calls a month, and a theme that shows up in 0.5% of them — a component that fails, a competitor being mentioned, confusion about a new plan. That is about 195 mentions a month. Enough to act on.
Now sample 2.5% of the calls, which is generous by industry standards. Those 195 mentions become about five calls. Five is indistinguishable from noise. It does not clear the bar of anybody's attention, and it will not survive a meeting where someone asks whether it is a real pattern.
Nothing was wrong with the sample. It is doing exactly what a sample does: estimating things that are frequent. The problem is that the valuable finding is not frequent — it is early.
These are illustrative figures, chosen round to make the arithmetic visible. Substitute your own volume and the conclusion holds: the rarer the theme, the more completely sampling erases it.
The three categories that eat the ranking
There is a second reason theme-finding fails on samples, and it survives even at full coverage if you rank by volume.
The reason people call is operational and stable. Billing, delivery, a password. Three or four categories take most of the volume and they do not move month to month. A dashboard sorted by size returns those categories every week, correctly, until whoever asked for it stops opening it.
What is worth knowing is what changed. A theme that was 0.2% last month and is 0.9% this month matters far more than the category that has been 30% for two years. So the ranking has to be by movement, not by mass — and movement can only be computed against a stable baseline, which means the categories have to be the same ones as last month.
What the customer said in passing
The most valuable material in a contact centre corpus is not the reason for the call. It is the sentence next to it.
Someone rings about an invoice and, while the agent is looking it up, mentions that they tried the new product and something about it did not work. They did not file a complaint. They did not answer a survey — they were not sent one. They did not write a review, because a person who is mildly disappointed rarely does.
That sentence exists in exactly one place: the recording. And it is invisible to any process that classifies the call by its reason, because by that measure this was a billing call.
What this means for how you buy
If you are evaluating call analysis, the question is not how accurate the scoring is. It is which of the two jobs the system is built for, because the two have opposite requirements.
Rating agents needs a defensible sample, a rubric, and evidence a supervisor can check. Finding themes needs the census, a taxonomy that stays stable so change is measurable, and analysis of the customer's turns rather than the agent's script — which is repetitive by design and drowns everything else.
A system built for the first can be pointed at the second, and it will return something. It will return the three categories you already knew about.