✦ Awe Reads. We write to help you see things a little clearer.

An older woman sits on the edge of a bed at night pressing a phone to her ear, her face lit by the screenAI-generated image
Technology

How Many Seconds Does It Take to Clone a Voice? Just 3

12 min08/14/2026

Three seconds of audio and a voice is cloned. The worse news: people got worse at recognising real speech, not fakes.

50 150

It is 11:40 at night. Unknown number. You pick up and hear your daughter crying. There has been an accident, police are there, she needs money right now. It is her voice. Her rhythm. Even the small catch of breath before she says "Mom."

It is not her.

Cloning a voice takes three seconds of recording. Telling the result from the real thing by ear is close to impossible, and that has now been measured: in a 2026 study, 1,768 people listened to clips and caught synthetic speech 71-73% of the time. A software detector in the same experiment was right 94.5% of the time. Your ear loses to the machine by more than twenty percentage points.

But the headline number in that study is a different one. Over five years people did not get better at catching fakes. They got worse at recognising real speech: 64.1% correct, down from 72.7% in 2021. A phone call is ceasing to prove anything at all, and that holds even when nobody on the line is trying to defraud you.

So what protects you is not sharper hearing. It is a habit: hang up and call back yourself. Below: how many seconds of audio a scammer needs and where they get it, why "listen for robotic pauses" is bad advice, and the one question that breaks any clone. That last one comes near the end.

How many seconds of audio does it take to clone a voice?

Three seconds. Not three minutes, not an hour of conversation. Three seconds of ordinary speech.

In 2023, Microsoft researchers published VALL-E, a system that can produce "high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker." The technical term is zero-shot synthesis (the model handles a person it has never heard before, with no separate training on that person). Before this, a convincing clone needed tens of minutes of clean studio audio and an engineer to supervise it.

A separate test backs the number up. McAfee engineers took free online tools and got an 85% voice match from a three-to-four-second clip. With more time spent tuning the clone, the match reached roughly 95%. The same report surveyed 7,054 people across seven countries: one adult in four had already run into a voice scam of some kind, and 77% of those who fell for one lost money.

So the only remaining question for a scammer is where those three seconds come from. The answer is uncomfortable.

Where your voice is sitting right now

  • Stories and short video. Any clip where you talk to camera. Five seconds is plenty.
  • Voice notes. A forwarded voice message keeps travelling through other people's chats, and you stop controlling it the moment you send it.
  • Work calls and webinars. Meeting recordings often sit behind an unprotected link.
  • Your voicemail greeting. Twelve words in your own voice, available to anyone who dials the number.
  • A normal incoming call. Saying "hello, who is this?" is already enough material.

The US Federal Trade Commission puts it plainly: a criminal needs "a short audio clip of your family member's voice - which he could get from content posted online - and a voice-cloning program." No phone access. No hacking. No leaked database.

Fine, a voice is cheap and easy to clone. But surely you would recognise your own child. That confidence is the part worth examining.

Can you tell an AI voice from a real one by ear?

Barely. People are poor at spotting synthetic speech in general, and practice hardly moves the needle. This is not about careless listeners. It is about everyone.

Researchers at University College London ran the test on 529 participants. People heard clips in English and Mandarin and judged each as real or generated. They identified the fakes 73% of the time. Some participants were then given examples of deepfakes (fakes assembled by a neural network: software that learns from examples and then produces something similar on its own) to train their ear. Accuracy improved by 3.84 percentage points. Statistically real, practically useless: being wrong a quarter of the time instead of slightly more than a quarter is not a defence.

The reason sits in how voice recognition works. Your brain does not store a full recording of a person. It keeps a handful of coarse anchors: pitch, pace, the shape of a familiar phrase, one or two pet words. Everything else it fills in, the way it fills in a face from a silhouette in a dark room. Synthesis copies exactly those anchors, because those anchors carry the most speaker information. So a clone does not slip past your recognition system. It lands right in the middle of it.

Then add the phone. The line clips the frequency range, compression eats the fine texture, and traffic or a stairwell masks whatever survives. The small synthesis artefacts you might catch on headphones in a quiet room never reach your ear on a real call.

The limits are worth stating honestly. That experiment used a single female speaker, and the share of fakes in the sample was higher than it would be in real life. The authors also noted they used an older synthesis method, not the state of the art. In other words, real conditions are probably worse, not better.

Three years later the test was repeated on a much larger group. The headline result had nothing to do with fakes.

What changed by 2026: why do people stop trusting real voices?

People did not get worse at catching fakes. They got worse at recognising the truth.

Researchers at Fraunhofer AISEC (a German applied research institute working on cybersecurity) collected 1,768 participants and 35,532 judgements of clips produced by 138 different synthesis systems, then compared the results against a matching study from 2021:

  • Accuracy on synthetic speech: 72.9% in 2021, 71.2% in 2026. Essentially flat.
  • Accuracy on genuine speech: 72.7% in 2021, 64.1% in 2026. A drop of 8.6 percentage points.

The authors call it a skepticism shift. Put simply: after years of deepfake headlines, people started suspecting fakery where there was none. A real human speaks into the phone and gets filed as a robot.

That reframes the whole problem. The usual warning is "you will be fooled by a fake voice." The wider cost is that a phone call stops proving anything at all, and that spreads well past the calls criminals actually make. A mother doubts her son. A court doubts a recording.

There is also an awkward practical conclusion for most articles on this topic. Checklists like "five signs of an AI voice: metallic tone, odd pauses, breathing that is too even" train precisely the skill that declined in the study. People start straining to hear, and they miss in both directions. Machines, for what it is worth, do fine: the detector in that same experiment scored 94.5%. You just do not have one in your handset.

So the question changes shape. Not "how do I hear it," but "how do I verify without relying on my ears." Before the answer, look at how these calls are built. The script is far easier to spot than the timbre.

How do AI voice scam calls work? Three scripts

A cloned voice almost never works alone. It is dropped into a ready-made pressure routine, and there are mainly three.

1. The relative in trouble

Decades old, now with a real voice attached. A crash, an arrest, a hospital, "please don't tell Dad." The tell is not the sound, it is the shape of the conversation: urgency, secrecy, and a ban on checking. You are asked to stay on the line, tell nobody, and settle it within fifteen minutes.

The FTC describes the payment stage this way: "If the caller says to wire money, send cryptocurrency, or buy gift cards and give them the card numbers and PINs, those could be signs of a scam." The logic is simple. Those transfers are nearly impossible to reverse.

2. The boss's voice

This one uses the org chart. A calm, familiar executive voice asks you to pay an invoice today, send a document, or confirm new bank details. Arguing with a director feels rude, and double-checking feels ruder. Kaspersky describes the mechanism plainly: the level of authority a caller appears to carry raises the odds that people comply without stopping to think.

3. The bank or the police

Globally the biggest earner. In Singapore, where police publish a detailed breakdown, government official impersonation ranked second by total money lost across all scam types in 2025. In the United States, people lost nearly a billion dollars that year to callers posing as businesses, and impersonated banks accounted for the largest share of that. Another $920 million went to callers posing as government agencies.

In practice it feels mundane. First the "bank security team" rings: a suspicious charge has been flagged, please confirm it was not you. You say it was not. The call is then handed politely to an "officer" who already knows your name and the last digits of your card. You are asked to keep it quiet, because an investigation is underway, and to move your money to a safe account. The voice is calm, the phrasing official, and there is office noise in the background. None of that appears on any list of AI voice red flags, because what has been faked here is not the sound but the situation.

What all three share: you are hurried, and you are steered away from checking. The voice exists only to stop you doubting long enough. National reports show what that costs.

How much money do people lose to voice scams?

Losses are rising almost everywhere, and the "almost" is the useful part. All figures are for 2025, each from a national regulator or police force:

  • United States: $3.5 billion lost to impersonation scams. Nearly one in three fraud reports fell into that category, and losses have roughly tripled since 2020. Total reported fraud losses hit $16 billion, about 25% higher than the year before (FTC, 2026).
  • Australia: AU$2.18 billion, up 7.8% year on year. The National Anti-Scam Centre specifically flags "increasing sophistication in scam activity through Artificial Intelligence."
  • Singapore: S$913.1 million, down from S$1.1 billion a year earlier, and this one is a fall: cases dropped by almost a quarter, 24.8%. Police prevented roughly S$348 million in further losses and seized about S$140 million more.
  • Russia and the CIS: 29.3 billion roubles taken from accounts in 2025, 6.4% more than the year before, with the number of unauthorised transactions up 31.2%. Thefts became more frequent but smaller. Banks blocked 13.9 trillion roubles' worth of attempts, but refunded victims only 1.7 billion, 5.9% of what was taken, down from 9.9% a year earlier (Bank of Russia data as reported by Interfax, February 2026).

Singapore deserves a second look. It shows the curve does not have to go up. What worked there was not better listening: it was payment holds, mandatory transfer delays, and relentless public warnings. The technology is identical in every country. The difference is the procedure wrapped around the money.

The same logic scales down to one household. Do not try to outsmart the technology. Change the procedure.

Which anti-scam tips do not actually work?

Half the popular guidance on this topic is useless, and some of it actively backfires. Worth clearing out before the part that works.

  • "Listen for unnatural pauses and breathing." That is precisely the skill the data shows declining. Worse, trained suspicion misfires onto real people: correct identification of genuine speech has fallen to 64.1%.
  • "Call that number back." Numbers are spoofed (caller ID spoofing: inserting someone else's number into the display, a routine technical trick that requires hacking nothing). Calling back reaches the same person who just lied to you.
  • "Ask for a code word during the call." Only works if the word was agreed in advance. Inventing one mid-call is pointless: you end up dictating to the caller the exact thing you then ask them to repeat.
  • "Ask a personal question." Personal and private are not the same. Pet names, car model, home town, school: all of it sits in open profiles and takes about a minute to find.
  • "Install an AI voice detector." Detectors exist and they are good: 94.5% accuracy in lab conditions. But they analyse a file after the fact, not the live stream in your handset.

All five share one flaw. They try to win on the caller's territory, inside the call itself. You win outside it.

How do you verify the caller is really who they say?

The whole routine takes under a minute and never asks you to judge a voice. That is the point: it holds even if the clone is perfect.

Check 1. Hang up and call back yourself

This is the core rule, and it is the FTC's own wording: "Don't trust the voice. Call the person who supposedly contacted you and verify the story. Use a phone number you know is theirs."

The operative word is yourself. Not "call the number that rang you," not "dial the number they read out." Open your own contacts and dial from there. A criminal may keep the line busy, so if you cannot reach the person, the FTC suggests reaching them through another family member or a friend.

The obvious objection: what if it is a real emergency and I waste time? Thirty seconds has never changed the outcome of an actual accident. It changes the outcome of a fake one, because the whole scheme survives only until you verify.

Check 2. A family code word

One agreed word that only your household knows and that appears nowhere in your chats or profiles. Not the dog's name, not your mother's maiden name; both are findable in a minute. Something arbitrary: "apricot," "blue suitcase," "platform nine."

It works because a voice can be cloned but a memory cannot. A neural network reproduces timbre and delivery. It has no idea what you agreed over dinner in March.

One rule keeps it alive: never say it in ordinary conversation and never type it. It exists only for the frightening call.

Check 3. A question with no answer online

This is the question promised at the start. Ask something only the real person could know and that appears in no profile: what you ate last Sunday, the name of the neighbour downstairs, where the spare key lives, how the argument about the holiday ended.

The caller is now stuck. They are steering a voice in real time, they have no seconds to spare inventing a plausible detail, and any pause breaks the urgency the scheme runs on. The usual reaction is irritation and a snap back to the clock: "I don't have time, I'll explain later." That reaction is the answer.

One condition: the question must be private, not merely personal. "What's my dog's name" is weak, because the dog is all over your feed. "What did we forget to buy last Saturday" is strong.

What should you do if you already sent the money?

Move within the first few hours, because the odds fall sharply after that. In order:

  1. Call the bank and report a fraudulent transfer. By phone, using the number on the back of the card, not through in-app chat. Ask them to log the case and give you the reference number.
  2. File a police report. Without it, most banks will not open an investigation and payment networks will not attempt a recall.
  3. Save everything. The calling number, the exact time, screenshots of the transfer, a recording if you have one. That is your entire evidence base.
  4. Warn your circle. One stolen voice is usually worked through the whole contact list. If "your son" called you, the grandmother and the brother are next.

The odds are worth stating plainly: in Russia in 2025, banks refunded 5.9% of what was stolen, less than the year before. Even those refunds come from reports filed within hours. On the rest, nothing comes back.

How can you protect your voice from cloning?

Not completely, and pretending otherwise would be dishonest. But you can shrink the supply of usable recordings:

  • Lock down accounts where you regularly speak to camera, unless publicity is part of your job.
  • Drop the personal voicemail greeting. The default system message works just as well.
  • Do not answer unknown numbers with your voice. Silence cannot be sampled.
  • Check whether recordings of your work meetings are sitting behind an open link.

And an honest caveat: if you teach, host a podcast, sell through video, or simply live in voice notes, your voice is already out there and nothing will pull it back. That is why the three checks matter more than the hygiene. The goal is not to make yourself impossible to clone. The goal is to make the clone useless.

The short version

  • Three seconds of audio is all a scammer needs to clone a voice.
  • People spot a fake about 7 times out of 10, and training barely helps.
  • By 2026, people had got worse at recognising real speech: 64.1%, down from 72.7% five years earlier.
  • A voice is not proof. Proof is a call back, a code word, and a question from private memory.

Back to the call at 11:40 at night. You hear your daughter's voice and your stomach drops. That part is human and it is supposed to happen. Then you say one sentence: "I'm going to call you back." And you dial her number from your own contacts.

Today, send one meaningless word to the family chat and agree on it: if something is wrong, we say this word. Two minutes. A voice can be faked. A shared word cannot.

Sources

#how many seconds to clone a voice#clone a voice with ai#how to spot an ai voice scam#voice cloning scam#ai voice deepfake detection#family code word scam#fake voice call signs

Best ideas, in your inbox

Once a week, 3-5 articles actually worth reading. No spam, unsubscribe in one click.

Read this next

Frequently asked questions

Three seconds. Microsoft's VALL-E synthesises speech from a three-second sample, and McAfee engineers reached an 85% voice match from a three-to-four-second clip using free online tools.

Barely. In studies from 2023 and 2026, people identified synthetic speech 71-73% of the time, and training added under 4 percentage points. A software detector scored 94.5%.

Hang up and call back yourself, using a number from your own contacts rather than the one that rang you. If they do not answer, reach them through another relative or friend.

Anything absent from your chats and profiles, so not a pet's name or a mother's maiden name. An arbitrary phrase like blue suitcase works. Never say it in ordinary conversation.

Not entirely. You can reduce the supply: lock down video accounts, drop your personal voicemail greeting, and avoid speaking to unknown numbers. Verifying calls matters more than hygiene.

Comments

Sign in to leave a comment.

Sign in

No comments yet. Be the first!

Awe ReadsFollow Awe Reads
How Many Seconds Does It Take to Clone a Voice? Just 3