This blog summarises the key points from a recent article by Derek Corcoran at Scorebuddy explores how to build effective human oversight into AI QA scoring, including when scores should be reviewed and how to manage the process.
Each quarter, a survey of contact centre professionals, ran by ScorebuddyCX, takes place through a QA & CX Intelligence Pulse Report to understand how teams are using different approaches in practice.
The latest report, the AI Reality Check, includes responses from 600 professionals and found that AI evaluation is now common, with nine in ten contact centers using it to assess interactions.
But using AI to generate a score does not necessarily mean that score is checked by a person. When asked about human oversight, just 14% of contact centers said every AI-generated evaluation is reviewed.
For almost half of those using AI evaluation, human checks are carried out only sometimes, rarely or not at all. Others reported a more regular level of review, falling between these two extremes.
The findings suggest that human oversight is present in most AI-enabled QA environments, but it is not always embedded into the day-to-day process. In some contact centers, reviewing AI scores appears to be a consistent responsibility, while in others it may depend on whether there is enough time to do it.
Having someone manually validate every score produced by AI would largely defeat the purpose of using AI for QA in the first place. AI can evaluate up to 100% of customer interactions, while a QA team simply cannot review that volume manually.
Re-checking every result would therefore recreate much of the workload that AI was introduced to reduce and undermine the value of having broader coverage.
However, full coverage does not mean AI can be left to operate without oversight. An effective AI-based QA process still needs the right configuration, documentation, training and ongoing adjustments.
Human involvement remains part of that process once the technology is in place, particularly as scorecards change and customer contact evolves over time.
Our interpretation, rather than something directly measured by the survey, is that reviews are less likely to happen when nobody has clear responsibility for them.
When contact volumes increase and teams become busier, an activity that is not built into the regular workload can easily be pushed aside.
That matters because an inaccurate AI score can have consequences beyond a dashboard. It could be discussed during an agent’s one-to-one, contribute to a performance review, or point to a scorecard question that the AI is repeatedly interpreting incorrectly. These are among the issues that can contribute to AI QA failing in contact centers.
Regular human checks can therefore help maintain the accuracy of AI scoring. But accurate scores do not automatically mean agents will accept or trust them — which is a separate issue explored in our article on whether agents trust AI QA scores.
Human review does not need to turn AI QA into a manual exercise. A workable approach can be built around four areas: deciding what requires review, assigning responsibility, checking AI against human evaluations, and keeping a record of challenges.
Not every AI-generated result needs to be checked before it is used. Identify the scores where human oversight matters most, taking into account your business objectives, regulatory obligations, team structure and the locations in which you operate.
Scores used in agent performance reviews are a logical place to start, as are compliance failures and any results that an agent has challenged.
Where checking every score is unnecessary, decide how often samples should be reviewed and who owns that task.
A named QA lead with a weekly review built into their responsibilities is easier to maintain when workloads increase than an open-ended arrangement where someone checks scores whenever they have spare time.
Take a sample of interactions and have both the AI and your most experienced evaluators score them.
Comparing the results question by question can highlight where the AI is interpreting a scorecard differently from your human team. Those questions can then be clarified or adjusted.
Even with the 90%+ accuracy reported for AI Auto Scoring, some results will still differ from a human assessment, making regular calibration useful.
Agents should have a straightforward way to request a review when they believe an AI score is incorrect.
Record what was reviewed and whether the score was changed. Over time, this creates a useful record of recurring issues, showing which scorecard questions are most often corrected and whether human review is being focused on the areas where errors occur most frequently.
It also provides a documented process for explaining how an individual score was reached if it is later questioned.
Human review should not be viewed as something that is only needed while an AI QA system is being introduced. It is not simply a temporary measure used to calibrate the technology before moving towards full automation.
The reason is that AI-generated scores can influence what happens next. They may shape coaching conversations and affect how an agent approaches their next customer interaction. At the wider business level, patterns identified through QA can also feed into decisions and insights outside the contact center.
With those outcomes depending on the accuracy of the underlying scores, having a process for human oversight remains an important part of making AI-generated QA results useful.
This post has been re-published by kind permission of ScorebuddyCX - view the original article.
Reviewed by: Robyn Coppell