Deepfake
In social engineering, a deepfake is synthetic audio, video or imagery generated to impersonate a specific real person convincingly. For security purposes the important case is not the fabricated video: it is cloned voice on a telephone call, because that is the channel most organisations still treat as proof of identity.
How it works
Generative models trained on a person’s recorded speech can reproduce their voice from a short sample, and public sources supply the sample readily: conference talks, interviews, podcasts, webinars, voicemail greetings. Real-time synthesis makes an interactive call possible, so the caller can respond to questions rather than play a recording.
The attack that follows is not new. It is vishing and business email compromise with the weakest part of the story repaired. Previously an urgent request from a senior figure arrived by email and the obvious control was to telephone and confirm. The clone attacks the confirmation step itself.
What goes wrong
Verification procedures assume the voice is the evidence. Payment authorisation, credential resets at the help desk, and approvals for out-of-band requests are all built on somebody recognising who they are speaking to, and that assumption is now unsafe for anyone whose voice is publicly available, which is every executive and most managers.
Training makes it worse when it teaches people to detect fakes. Detection advice ages faster than the models, and an employee who has been trained to look for artefacts will be confident in the wrong direction. The control that survives is procedural rather than perceptual: verification through a separate channel the requester did not choose, a call back to a number from the directory, dual authorisation above a threshold, and an explicit rule that urgency is a reason for more checks rather than fewer.
The other change is scale. Personalised, well-informed pretexts used to be expensive to produce and were reserved for high-value targets. They are now cheap, so spear phishing quality reaches ordinary volume campaigns, and advice built on spotting poor language no longer applies.
Where this shows up in an audit
Simulation exercises are where this is tested, and the scope has to move with the threat: an exercise limited to email measures a control that attackers have already routed around. Voice-based scenarios are run under explicit written agreement, with named participants, agreed limits and a debrief, because the ethical boundaries are tighter than in an email exercise. What we measure is whether the procedure held, not whether the individual was fooled, since the individual is not the control. Note that the EU AI Act introduces transparency duties for synthetic content; confirm the applicable provisions and dates before relying on them. This is part of how we rehearse voice and message-based social engineering.