Your CFO's voice is now an attack surface
What AI voice cloning means for your business, and the controls that actually stop it. The fix is mostly process, not product.
Your accounts payable clerk gets a call at 4:40pm on a Friday. It’s the CFO. The voice is right: the accent, the clipped delivery, the way he trails off at the end of sentences. He’s stuck between flights, a supplier payment was rejected, updated bank details are in the email he just sent, and it needs to go out before the cut-off.
Everything about that call is fake except the voice, and the voice was cloned from a webinar recording sitting on your company’s YouTube channel.
We’re writing about this because it has moved from a conference-keynote curiosity to something we now see in client environments. If you approve payments, run a help desk, or have executives who speak publicly, this is a live operational risk, and unlike most security threats, the fix is mostly process rather than product.
What voice cloning actually is
Voice cloning uses machine learning to build a synthetic model of a specific person’s voice. The system analyzes recordings and extracts a vocal fingerprint: pitch range, cadence, accent, breath timing, the small idiosyncrasies that make someone recognizable. That model then generates speech from typed text. The person says things they never said, in a voice their own colleagues can’t distinguish from the real thing.
The mechanics matter less than the economics. Three things have changed.
- Sample requirements collapsed. Older systems needed hours of clean studio audio. Current pre-trained algorithms can recreate a person’s voice from a three-second clip. Your CEO’s LinkedIn video is more than enough.
- Cost collapsed. Commercial platforms offer this at consumer subscription prices, with no technical skill required.
- Quality crossed the credibility threshold. The output no longer sounds like a robot reading a script. It sounds like a person on a bad mobile connection, which is exactly the pretext attackers use.
The barrier to impersonating your leadership team is now measured in dollars and minutes.
Where this hits the business
Payment fraud with a voice layer
Business email compromise is already the most financially destructive threat aimed at organizations. The FBI’s 2025 Internet Crime Report put BEC losses at just over $3 billion, and 86% of those losses moved by wire or ACH: fast, and usually gone before anyone notices.
Voice cloning doesn’t create a new category of fraud. It removes the last verification step your staff were relying on. The email looked slightly off, so they called to confirm, and now the call confirms it. In one documented 2019 case, the chief executive of a UK energy firm was talked into transferring hundreds of thousands of pounds to a fraudulent supplier account by a synthetic recording of his boss’s voice. That was seven years ago, with far cruder tooling.
The FBI now tracks this explicitly: AI-related complaints crossed 22,000 in 2025, with $893 million in associated losses.
We wrote up the email-only version of this attack, including how a mailbox gets opened weeks before any money is discussed, in the wire transfer that almost happened. Voice is the newer layer on the same script.
Help desk and identity verification
This is the one most businesses underestimate. If your service desk, internal or outsourced, will reset a password, re-enroll an MFA token, or add a device based on a caller sounding like the right person plus a couple of knowledge questions, you have a voice-authenticated administrative backdoor. Attackers know it, and cloned audio makes the “sounding right” part trivial.
Several of the largest breaches of the last few years began with a convincing phone call to a help desk, not an exploit.
Vendor and supplier impersonation
Bank detail change requests, invoice disputes, “our AP contact left, here’s the new one.” A follow-up call from someone who sounds like your long-standing contact at a supplier is persuasive precisely because the relationship is real and the voice is familiar.
Voice biometrics as an authentication factor
If any of your banking, payroll, or telecom relationships use voiceprint authentication, treat that factor as degraded. It was designed for an era when replicating a voice was hard.
Reputational and market risk
Fabricated audio of an executive discussing layoffs, an acquisition, or a quality problem can circulate widely before anyone verifies it. For regulated firms, listed companies, or anyone mid-transaction, the window between publication and correction is where the damage happens.
Why “train staff to spot it” is not a strategy
The instinct is to run awareness training on audio tells. The evidence says don’t rely on it.
UCL researchers played genuine and synthetic audio to 529 people and found detection unreliable: listeners correctly identified deepfakes only 73% of the time, with no difference between English and Mandarin, and showing examples of deepfakes beforehand improved results only slightly. Training lifted accuracy by less than four percentage points on average. That study used comparatively old algorithms; more recent work suggests the newest commercial systems are harder still, and that people are increasingly likely to distrust genuine audio as well. Accuracy on real samples fell from roughly 73% to 64%, while machine detectors held above 94%.
Read that as a design constraint. Roughly one in four synthetic calls gets through a human filter, and your own genuine calls will increasingly be doubted. Any control that depends on an employee’s ear is a control that fails a quarter of the time, on exactly the transactions attackers have chosen because they’re urgent and high-value.
The answer isn’t better listening. It’s removing voice from the trust chain entirely.
The controls we recommend
None of this is exotic. Most of it is policy and process, which is why it tends to get deferred, and why it’s cheap relative to a single successful wire fraud.
Make callback verification mandatory and directional. Any request to move money, change bank details, or alter payroll gets verified by calling back on a number already held in your master data. Never a number supplied in the request, and never by replying to the inbound channel. The caller does not get to choose the verification channel.
Enforce out-of-band, dual approval for financial changes. Vendor bank detail changes and payment instructions above a defined threshold require a second approver and confirmation on a separate channel. Document the threshold. Audit the exceptions.
Remove urgency as an override. Most successful voice fraud works because someone senior appeared to authorize skipping the process. State explicitly, from the executive team down, that no genuine request will ever require bypassing verification, and that nobody will face consequences for verifying. Executives who bristle at being verified are the vulnerability.
Harden help desk identity proofing. Replace voice-plus-knowledge verification with something an attacker can’t clone or look up: verification through an authenticated channel, manager confirmation for privileged resets, callback to the number of record, or an in-person or video step for high-risk accounts.
Agree on challenge phrases for the executive team. A pre-shared code word for urgent verbal requests involving money or credentials. Low-tech, effective, and free.
Audit your public voice footprint. Not to eliminate it, since public-facing leaders need to be public-facing, but so you know your exposure. Earnings calls, conference talks, podcasts, webinars, and promotional video are all viable training data. The right response is stronger process around the people most exposed, not silence.
Reduce reliance on voice biometrics. Where a banking or vendor relationship offers voiceprint as a factor, pair it with something else or turn it off.
Rehearse the incident path. Who does an employee call at 4:40pm on a Friday when they suspect a fraudulent request? Who contacts the bank to attempt a recall? At what point do you notify insurers, the board, law enforcement? Recovery of a fraudulent wire is measured in hours, and there is a real recall window if you move fast enough.
Check your insurance wording. Social engineering and funds-transfer fraud are frequently sub-limited or excluded from standard cyber policies, and losses where an employee authorized the transfer are precisely where the disputes happen. Read the policy before you need it, not after.
The legitimate side, and how to govern it
Voice synthesis isn’t going away, and much of it is valuable. Businesses use it to localize training and product content without re-recording, to produce consistent voiceover at volume, and to give branded assistants a recognizable identity. The most compelling use is medical: people facing conditions like ALS or throat cancer can bank their voice and keep speaking in it.
If your marketing, training, or customer experience teams are already experimenting, and in most organizations someone is, get ahead of it with a short policy:
- Written, revocable consent from anyone whose voice is cloned, including employees and executives
- Clear ownership of the voice model, and deletion on request or departure
- Disclosure to customers when a synthetic voice is used
- Approved platforms only, so voice models aren’t scattered across personal accounts
The legal picture is still forming, but existing frameworks already apply. Right of publicity and personality rights cover commercial use of someone’s voice. A growing number of US states have passed laws specifically targeting non-consensual synthetic likenesses. Defamation applies to fabricated statements, and fraud and impersonation statutes apply to the criminal use. In Europe, voice data can constitute biometric personal data under GDPR, which raises the consent bar considerably, and the EU AI Act imposes transparency obligations on synthetic media. If you operate across jurisdictions, treat consent and disclosure as the baseline regardless of where the lightest-touch rule sits.
What we’d do first
If you take three things from this:
- Rewrite your payment verification procedure this quarter so that no financial change can be authorized on the strength of a voice, and callbacks always go to a number of record.
- Fix help desk identity proofing, because it’s the highest-leverage and most commonly overlooked gap.
- Tell your people, in writing and from the top, that verifying is always acceptable, including when the voice on the phone belongs to the CEO.
The technology that makes this attack cheap isn’t going to get more expensive. But the attack only works against organizations where a familiar voice is sufficient authority to move money or grant access. That’s a decision your business gets to make, and it’s worth making deliberately rather than discovering after a wire has cleared.
If you’d like us to review your payment approval workflow, help desk verification standards, and executive exposure, get in touch. It’s a short engagement, and considerably cheaper than the alternative.