Contact center agents have always been trained to recognize suspicious behavior. An unusual request, inconsistent information, or a caller who seems to be following a script can all raise red flags.
But what happens when the person on the other end of the phone sounds completely natural?
Recent advances in voice AI are making that question increasingly important. Today's systems can carry on fluid conversations with realistic voices, natural pacing, and increasingly convincing emotional expression. The same technology creating exciting new customer experiences can also be used by bad actors.
New research suggests the threat isn't theoretical.
AI Is Changing the Economics of Voice Phishing
A July 2026 study examined whether current AI models could be used to automate voice phishing, or "vishing," attacks. Researchers tested several leading voice models in a large-scale experiment involving 4,100 U.S. adults and found that some AI voices were rated as comparably persuasive to human voices, with some even slightly outperforming the human baseline.
Across the five scam categories researchers tested, 16.5% of participants indicated that they would or might comply with the request. In "relative in distress" scenarios, that figure reached as high as 36%.
But perhaps the most important finding wasn't about persuasion at all. It was about economics.
Traditional voice phishing requires people to make calls, which limits how many victims an attacker can realistically target. The researchers concluded that human-operated vishing is generally uneconomical at U.S. wage levels, while AI-powered vishing appears economically viable using several of the models they tested.
AI doesn't necessarily have to become better at deception than humans.
It just has to make deception cheap enough to scale.
You May Not Be Able to Hear the Difference
There's another reason this matters for contact centers: simply listening for a "robotic" voice may no longer be an effective defense.
Separate 2026 research asked people to distinguish AI-generated voices from human-recorded voices in simulated vishing scenarios. Participants performed poorly, with an average accuracy of just 37.5%. Three-quarters of the AI-generated clips were classified as human by the majority of participants.
The characteristics people traditionally associate with authenticity, such as pauses, vocal variation, cadence, and emotion, can increasingly be reproduced by synthetic voices.
That changes the security equation.
If an agent can't reliably determine whether a caller is human based on how they sound, organizations need other signals to determine whether the conversation itself is suspicious.
Contact Centers Are an Attractive Target
Contact centers sit at an unusual intersection of customer service and security.
Agents are specifically trained to be helpful. They routinely help customers reset credentials, access accounts, update personal information, investigate transactions, and resolve unusual situations. In healthcare, financial services, insurance, travel, and other industries, they may also have access to extremely sensitive information.
Those responsibilities make the contact center a natural target for social engineering.
The risk becomes even greater if AI allows attackers to conduct thousands of convincing conversations at a fraction of the previous cost. An attacker doesn't need every call to succeed. At sufficient scale, even a relatively small success rate can become meaningful.
And because the caller may sound perfectly normal, putting the burden entirely on the individual agent becomes increasingly unrealistic.
The Conversation Itself Becomes a Security Signal
This is where conversation intelligence can play a much larger role.
Instead of asking an agent to independently recognize every potential threat, organizations can analyze conversations for patterns that might warrant additional attention. That could include unusual requests, repeated attempts to bypass standard procedures, unexpected combinations of account activity, or sudden increases in certain types of interactions.
Research published this summer in Scientific Reports explored a similar idea at the network level, using call-log patterns and machine learning to proactively identify voice-phishing networks rather than waiting until individual victims reported fraud.
For contact centers, the broader principle is important: security can't depend exclusively on whether one employee thinks a caller sounds suspicious.
The organization needs visibility across conversations.
Real-Time Guidance Can Become a Security Layer
Conversation monitoring becomes even more valuable when insights can reach the agent while the interaction is still happening.
Imagine an agent receives a seemingly ordinary call involving an unusual account request. As the conversation develops, real-time agent assist recognizes that certain behaviors require additional verification and reminds the agent of the appropriate procedure.
The agent doesn't need to diagnose whether they're speaking with a sophisticated AI system. They simply need to follow the right process.
This is where capabilities like accurate transcription, real-time agent assist, automated QA, and compliance monitoring can work together. MosaicVoice can help organizations identify important moments in conversations, provide timely guidance to agents, and evaluate interactions at scale rather than relying solely on manual call reviews.
The objective isn't to label every unusual conversation as fraud. It's to give organizations and their agents more visibility when something deserves a closer look.
Human Oversight Becomes More Important, Not Less
There's an interesting irony in all of this.
As AI becomes more sophisticated, human oversight may become more important.
Organizations will need people to investigate emerging patterns, refine verification processes, coach agents on new threats, and continuously update the guidance provided during conversations. Automated QA can help identify whether those processes are actually being followed across thousands of interactions rather than a small sample of calls.
The strongest defense will likely combine AI's ability to recognize patterns at scale with human judgment about what those patterns mean.
That's especially important because the threat itself will continue changing.
The Bottom Line
For years, voice phishing had a natural constraint: people.
Attackers needed someone to make the calls, carry the conversation, respond naturally, and persuade the person on the other end. Voice AI is beginning to remove that constraint.
The next generation of fraudulent calls may not sound robotic, awkward, or obviously suspicious. They may sound completely normal.
That means contact centers need to look beyond the sound of the caller and start paying closer attention to the behavior happening within and across conversations.
In the age of AI-generated voices, the question may no longer be, "Does this caller sound real?"
It may be, "Does this conversation look right?"