The 6 Steps to QA Support Tickets at Scale in 2026
Run the arithmetic on your own QA program before you trust its output. Most teams score one to two percent of tickets, which at typical team sizes works out to a handful of conversations per agent per month. At that sample size, the variance introduced by which tickets happened to get drawn is larger than the variance between your strongest and weakest agent. The scores are real numbers. They are not measuring what the dashboard says they measure.
The six steps to QA support tickets at scale are separate measurement from coaching, replace random sampling with risk weighting, auto-score the population and human-score the sample, calibrate reviewers before comparing them, pull coaching examples from the agent's own tickets, and verify the score moved. Coverage is not the first problem to solve. Sampling design is, because a bad sample scaled up is just a more expensive bad sample.
What a support QA program actually has to deliver
- Coverage of the population, not confidence in the sample. Two different jobs. Knowing your overall quality trend requires reading everything at some level. Knowing whether a specific agent improved requires reading a lot of that agent. One sampling strategy cannot serve both.
- Categories derived from the conversations. Risk weighting requires knowing what a conversation was about before you decide whether to review it. If your categories come from a manually maintained picklist, the themes worth reviewing most are the new ones nobody has a tag for yet.
- Account and revenue context on each interaction. A mediocre response to a trial user and a mediocre response to an account in renewal are the same score and different problems. QA programs that ignore this systematically under-review the interactions that cost the most.
- A separation between scoring and coaching. A number tells an agent where they rank. An example from their own queue tells them what to do differently. Most programs deliver the first and call it coaching.
The bottleneck is not review capacity. It is that random sampling is optimized for fairness to agents rather than for finding problems.
The 6 steps to QA support tickets at scale
1. Separate measurement from coaching
Decide which one each review serves, because they need opposite samples. Measurement wants a representative draw so the trend is unbiased. Coaching wants a deliberately biased draw toward an individual agent's weak dimension. Running one sample for both produces a trend nobody trusts and coaching nobody can act on.
Output: two review streams with different sampling rules.
2. Replace random sampling with risk weighting
Random sampling treats every conversation as equally worth reading, which is the one thing you know is false. Weight the sample toward interactions with a repeat contact, a low satisfaction score, an escalation, a long handle time, an account in renewal, or a theme that is new this week. You will read the same number of tickets and find several times more problems.
Output: a weighted queue rather than a random draw.
3. Auto-score the population, human-score the sample
Some dimensions are machine-checkable at full coverage: whether the response addressed the stated question, whether policy language appeared, whether the customer restated their question afterward. Others need a human: judgment, tone in a hard conversation, whether the right exception was made. Score the first category across every ticket and reserve human review for the second, which is what makes coverage and depth stop competing.
Output: full-population signal plus a small deep-review set.
4. Calibrate reviewers before comparing anyone
Two reviewers scoring the same conversation typically disagree more than two agents differ from each other. Until you have measured reviewer agreement on a shared set and closed the gap, agent-to-agent comparisons are measuring reviewer assignment. Recalibrate quarterly, because drift is continuous.
Output: a known inter-reviewer agreement rate, checked on a schedule.
5. Pull coaching examples from the agent's own tickets
Generic examples do not transfer. Once an agent's weakest dimension is known, search their own queue for two conversations that went badly on that dimension and one that went well, and coach on the contrast. The one that went well matters more than most managers expect, because it establishes the behavior is already available to them.
Output: three of the agent's own conversations per coaching session.
6. Verify the score moved
Treat every coaching conversation as a hypothesis with a follow-up date. Re-sample that agent on that dimension specifically four to six weeks later, weighted rather than randomly, so the follow-up is actually capable of detecting a change. Most QA programs skip this and therefore never learn which coaching works.
Output: a per-dimension before and after, per agent.
Random sampling is the wrong default
Worth stating directly, because it is the assumption everything else inherits. Random sampling exists in QA programs for a defensible reason: it feels fair, and it protects against managers cherry-picking evidence against people they dislike. Those are real concerns.
But the cost is that the conversations you most need to read are exactly the ones a random draw under-represents. Bad interactions are rare by definition. If four percent of conversations contain a real failure and you sample five per agent, you will usually find nothing, conclude quality is fine, and be wrong in a way no amount of additional random sampling fixes.
Risk weighting solves the detection problem and reintroduces the fairness problem, which is why step one exists. Use the weighted sample to find problems and a separate representative sample to report trends. Do not use the weighted sample to rank people.
This is the same instrumentation as identifying the primary driver behind a support contact, because both depend on knowing what a conversation was about before you decide what to do with it. For the tooling comparison, see CSAT tools for support QA and agent coaching and best customer support analytics tools.
How to build the weighted queue
The practical obstacle is that risk weighting needs signals your helpdesk does not compute: whether this conversation is part of a repeat contact, whether the theme is new, and what the account is worth.
Enterpret supplies those. Its adaptive taxonomy categorizes conversations by what the customer meant rather than by tag, which is what lets you weight toward themes that are new this week and to find every conversation touching an agent's weak dimension regardless of how customers phrased it. Its customer context graph attaches account, ARR, and renewal timing, so the queue surfaces the interactions where a mediocre response is expensive rather than merely imperfect.
Start with step two alone. Reweight next month's sample toward repeat contacts and low scores, keep the sample size identical, and compare how many real issues you find. If the count does not go up substantially, your sampling was not the bottleneck and you have learned something useful either way.
FAQ
How do I QA support tickets at scale?
Split the work by what needs a human. Auto-score the full population on checkable dimensions such as whether the stated question was addressed and whether the customer came back, then reserve human review for judgment, tone, and exception handling on a small deliberately weighted sample. Scaling human review linearly with ticket volume never works, and it is not necessary.
How do I score support quality without reading every ticket?
Score every ticket automatically on the dimensions a model can evaluate, and read a weighted sample for the rest. The important design choice is the weighting: pull toward repeat contacts, low satisfaction, escalations, new themes, and high-value accounts rather than drawing randomly, because failures are rare and a random draw of a few tickets per agent will usually miss them.
How many tickets should we review per agent?
Enough that the result is not dominated by which tickets got drawn, which for most teams is considerably more than the five or so they currently do. Rather than picking a number, check reviewer agreement first and then increase the sample until an agent's month-over-month score is stable when quality has not changed. If the score bounces without any underlying change, the sample is too small to act on.
How do I find coaching examples from an agent's own tickets?
Identify the agent's weakest scored dimension, then search their own conversations for two examples that went poorly on that dimension and one that went well. Coach on the contrast rather than on the score. Examples from someone else's queue are easy to dismiss, and the positive example is what proves the behavior is already within reach.
How does Enterpret support QA and agent coaching at scale?
Enterpret categorizes every support conversation with an adaptive taxonomy derived from the conversations themselves, which makes it possible to build a risk-weighted review queue and to find all conversations touching a specific issue or behavior regardless of wording. Its customer context graph attaches account, segment, and revenue to each interaction, so review effort concentrates where a poor response actually costs something rather than being spread evenly across a random draw.
If your QA sample is random, see how Enterpret's adaptive taxonomy lets you weight the queue by theme and risk instead.
Heading
Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.



