Threat Intelligence AI Fraud Prevention

Voice deepfakes and CEO fraud: how AI voice cloning is targeting payment authorisation

AI voice synthesis can replicate an executive's voice from a few minutes of public audio. Finance teams are taking calls that sound exactly like their CEO or CFO and authorising urgent wire transfers. Below: how the attack works, why it succeeds, and the controls that stop it.

11 August 2026
8 min read
Key takeaways
  • Voice synthesis tools can clone a voice from as little as three minutes of audio, using recordings from earnings calls, interviews, or YouTube videos
  • The attack pattern combines cloned voice with social engineering: urgency, secrecy, and authority bypass normal verification instincts
  • A callback to a known corporate number, not a number provided in the call, stops the fraud before money moves
  • Dual authorisation for all payments above a set threshold means no single person can move funds alone
  • The hardest training task is convincing staff that recognising the voice is no longer enough to trust a call

How the attack works

Business email compromise (BEC) fraud has relied on impersonation for years: a spoofed email from the "CEO" instructing the finance team to transfer funds. Voice deepfake attacks add AI-generated audio to that social engineering foundation, removing the last safeguard many finance teams trust: "I'd recognise his voice."

The technical barrier is low. Open-source voice synthesis models and commercial services can produce a convincing voice clone from three to five minutes of training audio. For executives at public companies, that audio is available through earnings calls, conference presentations, and media interviews. For executives at smaller businesses, LinkedIn video posts, recorded webinars, and company promotional videos provide the same raw material.

The attack follows a predictable script. The caller contacts a finance team member, poses as the CEO or CFO, and explains that a sensitive acquisition, regulatory matter, or supplier dispute requires an urgent payment today. The recipient is asked to bypass normal approval steps given the confidential nature of the matter. Authority, urgency, and secrecy work together to suppress the verification instinct.

Cases that show the pattern

In 2019, the CEO of a UK energy subsidiary took a phone call from someone he believed was the chief executive of the firm's German parent company. The caller's voice, accent, and speech patterns were convincing. He transferred approximately £200,000 to a supplier in Hungary within the hour. The money moved through a series of accounts and was never recovered. It was the first publicly reported case of AI voice cloning used in a BEC fraud.

In 2024, WPP's CEO Mark Read was impersonated in a Teams meeting using a WhatsApp voice clone and publicly available video footage. Attendees were asked to provide personal information and transfer funds. Suspicion grew, participants verified through a separate channel, and the attempt failed. Even so, the setup represented a significant escalation from phone-only attacks: a cloned voice combined with screen-shared video.

In the most financially damaging reported case, a finance worker at a multinational in Hong Kong joined a video conference in which the CFO and several colleagues appeared on screen and told him to authorise multiple transfers totalling approximately US$25 million. Every person on the call was a deepfake. The employee contacted head office through a separate channel days later, and that is when the fraud came to light.

Why finance teams are the specific target

Finance staff work in environments where speed, deference to authority, and confidentiality are normal professional values. A CEO asking them to handle something urgently and discreetly fits the pattern of routine executive communication. Fraudsters exploit professional norms, not naivety.

Red flags that precede payment fraud

Unexpected urgency, requests to bypass normal approval steps, instructions to keep the matter confidential, and requests to use a different payment method or account than usual appear consistently across documented cases. Any single one is unremarkable. Together, they are the pattern.

Callers also manufacture time pressure to block verification: "The transfer needs to complete before the markets close," "I'm in a meeting and need this done in 30 minutes," "I'll explain everything afterwards." Legitimate urgent transactions can survive a two-minute callback to a known number.

What to train your staff to look for

Train finance and operational staff on one core point: voice recognition is no longer a reliable authentication method. Knowing what to look for does not stop the fraud if the only available response is "this feels odd but the voice is right."

  • Unexpected contact channel. A CEO calling from a personal mobile, WhatsApp, or unfamiliar number is a signal, not proof, but enough to trigger callback verification.
  • Request to bypass normal process. Skipping dual authorisation, using a different account, or processing a payment without a purchase order or contract are all grounds to pause and verify.
  • Secrecy instruction. "Don't tell anyone about this yet" is one of the most consistent indicators of social engineering across fraud types.
  • Time pressure. Legitimate urgent payments survive a 90-second callback. Resistance to that delay is itself a signal.
  • Subtle audio anomalies. Current voice synthesis occasionally produces micro-pauses, unnatural rhythm, or background audio inconsistencies. These are not reliable on their own but can reinforce suspicion raised by other signals.

Controls that prevent payment fraud

Detection training matters but is not enough on its own. The controls that reliably prevent fraud are procedural, not perceptual.

Callback verification to a known corporate number

When a payment request arrives by any channel, the finance team member should end the call and dial the requester back on a number from the company directory or official records, never a number provided during the call or in an accompanying message. This single control would have prevented every documented voice deepfake payment fraud to date. Make it firm policy, not a guideline, and enforce it even when the requester pushes back on the delay.

Dual authorisation above a threshold

No single person should authorise a payment above a defined threshold without a second approver who verifies the request through a separate channel. Set the threshold low enough to catch most fraud attempts, which typically target amounts large enough to be worth the effort but small enough to avoid automatic review.

A code word for unusual requests

Some organisations give executives a verbal code to use when making requests that deviate from standard procedure. Staff agree on it in advance and never share it digitally. An executive who does not use it when asking for an exception signals that the call may not be genuine. The control needs almost no infrastructure, only advance agreement and a short briefing.

Payment process documentation

Recording the authorisation chain for every payment, including informal requests, creates a paper trail that makes fraud harder to conceal. Finance teams that can show who authorised what, through which channel, catch fraudulent requests before they complete.

Where AI makes this harder

Voice synthesis quality is improving faster than detection technology. Current AI audio detection tools carry meaningful false-positive rates in real-world conditions, making them unreliable as a frontline filter. Real-time deepfake detection integrated into phone infrastructure exists but has not reached wide enterprise deployment.

Procedural controls are the current best defence. Organisations that rely on technology to flag deepfakes before they reach staff are asking more than the technology can deliver. Organisations that have embedded verification procedures independent of detecting the fraud stop losses regardless of how convincing the synthetic voice becomes.

Tools that detect synthetic voice in real time

Several commercial platforms now offer detection built for voice deepfakes, with uptake growing fastest in financial services. Each takes a different technical approach and carries its own constraints.

  • Pindrop Pulse. Pindrop Pulse analyses 1,380 audio features per call, covering acoustic properties and liveness indicators. Major US banks use it at scale, and Pindrop reports accuracy above 99 percent on known synthetic voice models. The critical constraint is that the figure applies only to models Pindrop has already trained against. When an attacker uses a novel or custom-built voice cloning model, detection rates fall until Pindrop retrains on the new data. That lag can run to weeks or months.
  • Nuance Gatekeeper. Nuance Gatekeeper compares the caller's voice against a stored voiceprint for the claimed identity, rather than scanning the audio for synthesis artefacts. Lloyds Banking Group and HSBC both use it in contact centre operations. If an attacker calls claiming to be an internal executive and the voiceprint does not match, the system flags the call for review. The control works best when you have already registered voiceprints for your executives, which requires a pre-enrolment step.
  • Resemble Detect. Resemble Detect is an API-based classifier that operates in real time with latency under 100 milliseconds. Enterprise teams can integrate it into call workflows without a large telephony infrastructure change. Unlike the bank-focused platforms, any team can connect it through a standard API.
  • Microsoft Azure AI Content Safety (audio). Part of Azure's responsible AI suite, this tool includes audio analysis aimed at detecting synthetic speech. Microsoft has trained its models on English, which limits the tool's accuracy on calls in other languages or strong accents. Latency varies by region, which affects whether the tool can flag a call before a payment decision is made.
  • Human-in-the-loop verification. No detection platform on the market is reliable enough to serve as the sole control. A callback to a pre-registered number, kept independent of the channel through which the request arrived, stops the fraud regardless of how good the voice clone is. This procedural step remains the strongest single control regardless of which tool you deploy alongside it.

All current detection tools share one structural weakness: they depend on known voice models. Every vendor trains on synthetic audio it has collected or licensed. A model built to evade one of these systems will outperform it until the vendor retrains. Defence-in-depth, combining detection tooling with firm procedural controls, is the only architecture that holds.

Why the threat is accelerating

The technical barrier to voice cloning has dropped faster than most security teams expect. Three years ago, producing a convincing clone required substantial compute resource and hours of training audio. Today, platforms such as ElevenLabs and Descript generate a working voice clone from 30 to 60 seconds of audio. Earnings calls and LinkedIn video posts give attackers enough raw material without contacting the target.

Cost is no longer a meaningful barrier. Commercial API services generate a convincing clone for under £10. That makes mass targeting viable: an attacker can build voice clones of executives across dozens of organisations and run campaigns at scale.

Real-time voice conversion tools remove the need for pre-recorded audio. Open-source variants on HuggingFace let an attacker speak into a microphone while software converts the output to match the target's voice. You cannot spot the short processing delay in a normal call. The attacker needs nothing beyond free software and a microphone.

Multilingual models have extended the threat to cross-border fraud. Current systems clone accent and intonation, not just phoneme patterns, which means an attacker no longer needs to speak the target's language to produce a convincing impersonation. International CEO fraud operations are easier to mount as a result.

Call timing is part of the attack. Social engineering campaigns target Monday mornings and financial quarter-end periods, when time pressure on payment approvals is part of the working day. Voice deepfake attackers follow the same pattern because urgency in the environment makes verification feel like an obstacle.

UK Finance data puts the financial stakes in context. In H1 2024, authorised push payment fraud cost UK victims £213.7 million. Voice-initiated fraud is a subset of that figure, and its share is growing as the tools to generate it become cheaper and more accessible.

For a broader look at how deepfake technology is affecting identity verification, see our guide on deepfakes and identity fraud.

Ryland Deakin
About the author
Lead Consultant, Cyvra · CISM · CompTIA Security+ · MCP

Ryland has delivered cybersecurity, compliance, and IT management programmes for regulated organisations across the UK and the Netherlands for over 20 years, including senior roles at Microsoft, ING, IPsoft, PPHE and more. View full profile

Talk to Cyvra

Does your team know what to do when the CEO calls?

We help organisations build the payment verification controls and staff training that prevent fraud before it happens.