DEV Community

Cover image for That Voice May Not Be Real
tyjean
tyjean

Posted on

That Voice May Not Be Real

For years, voice felt like identity.

If your boss called, you listened.

If your child sounded panicked, you reacted.

If a senior executive appeared in a video meeting, you assumed the meeting was real.

That assumption is now dangerous.

AI voice cloning and deepfake video tools have changed the trust equation. A scammer no longer needs to sound like a professional actor. They may only need a short audio clip from a podcast, a webinar, a social media video, or an old voice message. With enough pressure and timing, a fake voice can make a payment request feel urgent. A fake video call can make an internal approval feel legitimate.

The most dangerous deepfake does not need to fool everyone forever. It only needs to fool one person for a few minutes.

A few minutes is enough to send money, approve a vendor change, reveal a verification code, share internal documents, or ignore a normal security process.

So the question is no longer:

Does it sound like them?

The better question is:

Can this request be verified outside the channel where it arrived?

Voice Is No Longer Proof of Identity

Voice used to be a strong social signal. We recognize family members by tone, executives by speaking style, and colleagues by rhythm and phrasing.

Voice cloning attacks exploit exactly that trust.

The U.S. Federal Trade Commission has warned that scammers can use AI to enhance family emergency scams. The setup is familiar: a panicked voice claims to be a child, grandchild, spouse, or friend in trouble, and the victim is pressured to send money quickly. The difference is that the voice may now sound much more convincing.

The same pattern applies in business.

A finance employee may receive a call from someone who sounds like the CEO. A vendor may send a voice note that sounds like an account manager. A manager may join a video meeting that appears to include familiar executives.

The danger is not only the synthetic voice. It is the social pressure around it.

Urgency Is the Attack

Most voice cloning and deepfake scams are not built around long conversations. They are built around pressure.

Common patterns include:

  • "I need this wire transfer approved now."
  • "Do not discuss this with anyone yet."
  • "This is confidential."
  • "I am in trouble and need money immediately."
  • "The account is locked, and you must verify it now."
  • "This vendor payment must go out before close of business."

The voice or video lowers your guard. The urgency makes you move before you verify.

That is why the safest response is procedural, not emotional.

If money, access, credentials, sensitive data, or authorization is involved, stop the interaction and verify through a separate channel.

Deepfake Video Calls Are Not Science Fiction

Many people understand that a voice can be faked, but still trust video calls.

That confidence is also outdated.

One of the most widely reported cases involved a finance worker in Hong Kong who was tricked into transferring roughly $25 million after joining a video call with what appeared to be the company’s CFO and other colleagues. Reports later described the participants as deepfake impersonations.

The key lesson is not that every video call is fake. The lesson is that video presence alone should not override payment controls.

If a video meeting asks someone to move money, change bank details, reveal credentials, bypass procurement, or ignore normal approval steps, the video is not enough.

The Verification Rule: Leave the Original Channel

The simplest defense is also the strongest:

Do not verify a suspicious request inside the same channel that delivered it.

If the request came by phone, do not trust the caller ID or the number they give you. Call back using a number already saved in your contacts or listed in an official internal directory.

If the request came through a video meeting, do not ask the people inside the meeting to prove they are real. End the meeting and confirm through established company channels.

If the request came through a voice message from a family member, call that person using a known number. If they do not answer, contact another trusted person.

If the request came from a vendor, verify through the vendor’s previously established contact and your internal approval process.

Scammers want to keep you inside their controlled environment. Verification means leaving it.

Preserve the Audio or Video Before It Disappears

If you receive a suspicious voice message, video file, meeting recording, or audio clip, preserve it as close to the original form as possible.

Do not trim it.

Do not convert it.

Do not add subtitles.

Do not screen-record a screen recording if you can download the original.

Do not forward it through apps that compress media unless you have no other option.

Also save the surrounding context:

  • caller ID or account handle
  • chat history
  • meeting invitation
  • requested payment details
  • timestamps
  • email headers, if relevant
  • screenshots of the request

This matters because an audio or video file is not just content. It is evidence of what was submitted to you. If the case later goes to a fraud team, platform, bank, insurer, employer, or law enforcement, the original file and context are much more useful than a vague description.

Use Technical Analysis as an Early Warning Layer

Human perception is not enough. But a single AI detector should not be treated as a final judge either.

A better workflow combines:

  • original file preservation
  • audio or video analysis
  • file fingerprinting
  • context review
  • independent identity confirmation
  • a written report that can be shared

ShanHaiYin supports AI content identification and provenance verification across images, text, audio, and video. For suspicious audio or video files, it can help analyze whether the file shows signs of AI synthesis, deepfake manipulation, abnormal editing, or other technical risk indicators. It also generates a report with the main conclusion, reasoning, request number, and SHA-256 file fingerprint.

That report should not be described as a court ruling. It is not a replacement for law enforcement, platform review, or human investigation.

Its practical value is different:

It helps you move from:

"I feel like this voice is suspicious."

to:

"We preserved the submitted audio/video file, ran a technical verification, and the report found indicators that require manual review before any payment or authorization."

That shift matters in companies, families, banks, marketplaces, and support workflows.

A Business Workflow for Deepfake Requests

Every organization should assume that voice and video impersonation will eventually target someone inside the company.

The defense does not have to be complicated.

Start with these rules:

  1. No payment is approved by voice alone.
  2. No bank-account change is accepted from a call or video meeting alone.
  3. Any urgent executive request must be confirmed through a second channel.
  4. Large transfers require two-person approval.
  5. Confidential projects still follow payment controls.
  6. Finance staff have explicit permission to pause suspicious requests.
  7. Suspicious audio and video files are preserved and analyzed.

The most important cultural rule is this:

Slowing down an unusual payment is not disobedience. It is security.

Deepfake scams work best when employees feel they are not allowed to question authority. Good controls remove that pressure.

A Family Workflow for Voice Cloning Scams

Families need a simpler version.

The FTC has warned about AI-enhanced family emergency scams, and many security professionals now recommend a family safe word or code phrase.

A practical family rule could be:

  1. If a call asks for urgent money, hang up and call back using a known number.
  2. Ask for a family code word that is never posted online.
  3. If someone says "do not tell anyone," tell another trusted family member immediately.
  4. Do not send money through gift cards, crypto, wire transfer, or unfamiliar payment links under pressure.
  5. Save suspicious voice messages or videos for review.

The point is not to make family life paranoid. It is to make one calm rule stronger than one emotional moment.

What to Do If You Already Received a Suspicious File

If you have already received a suspicious audio or video file, do this:

  1. Stop responding inside the same channel.
  2. Save the original file and surrounding messages.
  3. Upload the file to an audio/video verification tool such as ShanHaiYin.
  4. Download the report.
  5. Contact the real person through a known number or internal directory.
  6. If money or credentials were involved, notify your bank, employer, platform, or local authorities as appropriate.

Do not argue with the scammer.

Do not accuse them inside the call.

Do not keep giving them more time to pressure you.

Your goal is not to win the conversation. Your goal is to leave the trap.

Copy-Paste Responses

For a suspicious executive request:

This request involves company funds or authorization. I need to verify it through our established approval process. I cannot approve payment based only on a call, voice message, or video meeting.

For an internal finance alert:

We received a voice/video request involving payment or authorization. The media file has been preserved and submitted for AI synthesis/deepfake risk analysis. Please pause processing until identity is confirmed through an independent company channel.

For a family emergency call:

I am going to hang up and call you back on your usual number. If this is real, we will still handle it. I will not send money during this call.

For reporting the incident:

We received an audio/video file from an account claiming to be [identity]. The sender requested [payment/access/credentials/authorization]. We preserved the original file, generated a technical verification report, and are submitting the file, report, and related messages for review.

The New Rule: Verify the Person, Not the Performance

AI voice cloning and deepfake video attacks are powerful because they attack a very human habit: trusting familiar voices and faces.

But a voice is no longer enough.

A face on a screen is no longer enough.

A video meeting is no longer enough.

For any request involving money, access, confidential information, credentials, or authorization, the standard must change.

Do not ask:

Does this sound like them?

Ask:

Have we verified this through a trusted path outside the call?

If you receive suspicious audio or video, preserve the original file, run a technical review with a tool such as ShanHaiYin, and confirm identity through a known number, internal directory, family code word, or written approval process.

The future of trust will not be built on recognizing voices.

It will be built on verification.

References

  1. FTC: Scammers use AI to enhance their family emergency schemes https://consumer.ftc.gov/consumer-alerts/2023/03/scammers-use-ai-enhance-their-family-emergency-schemes
  2. FTC: Approaches to Address AI-enabled Voice Cloning https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/04/approaches-address-ai-enabled-voice-cloning
  3. FTC: New protections to combat AI impersonation of individuals https://www.ftc.gov/news-events/news/press-releases/2024/02/ftc-proposes-new-protections-combat-ai-impersonation-individuals
  4. The Guardian: Hong Kong company deepfake video conference call scam https://www.theguardian.com/world/2024/feb/05/hong-kong-company-deepfake-video-conference-call-scam
  5. National Cybersecurity Alliance: Safe Word / AI Fools https://www.staysafeonline.org/campaigns/safeword

Top comments (0)