Why does voice cloning make voice biometrics unreliable as a sole authentication factor?
Modern AI can clone a target's voice from just 20–30 seconds of audio and bypassed voice-ID on all four major assistants with ~95% success — cheaply and easily, so voice alone isn't enough.
Voice cloning uses AI to generate synthetic voices indistinguishable from real people. The process needs only 20–30 seconds of a target's audio to train a cloning model (using tools like Descript Overdub, Resemble.ai, or ElevenLabs), then synthesises commands to test voice-ID-protected systems.
Results: minimal recording time of 20s suffices for high-quality cloning; voice-ID systems were bypassed with a 95% success rate across all 4 systems (Alexa, Bixby, Google Assistant, Siri).
The technology is easily accessible, cheap, and highly effective — voice biometrics as a sole authentication factor is inadequate. Real threat: in 2019 the CEO of a UK energy firm was defrauded of €243,000 by a cloned voice imitating the parent-company boss.
Tip: The €243k CEO-fraud case shows the threat isn't just unlocking gadgets — voice cloning enables social-engineering fraud against humans, who are even easier to fool than a device.
Go deeper:
Audio deepfake (Wikipedia) — voice cloning and its threat to bank/voice-ID systems.