Back

AI Audio Watermarking Explained: How to Check If a Voice Was Made with AI

Shield with a check mark over a sound wave, illustrating audio watermark detection

As AI voices become indistinguishable from real recordings, people need a reliable way to tell them apart. Listening is no longer enough. Watermarking puts the answer inside the audio itself - inaudible to people, readable by software.

This article explains how audio watermarking works, exactly how Pynokio marks the voices it generates, and how anyone can check a file.

What is an audio watermark?

An audio watermark is a tiny, structured pattern embedded into the sound wave. It is designed to be imperceptible to listeners and to survive everyday handling of the file: renaming, re-saving, converting to MP3, trimming or adding background music. A matching detector analyses a recording and reports whether the pattern - and the message it carries - is present.

Unlike metadata, a watermark lives in the audio signal itself. Stripping file headers or uploading to a platform that removes metadata does not remove it.

How Pynokio marks generated voices

Pynokio uses three complementary layers:

1. An inaudible watermark in the sound

Every generated speech file - text to speech, voice clone previews and voice design previews - carries a watermark created with AudioSeal, an open-source audio watermarking technology from Meta. The watermark holds a 16-bit identifier that belongs to Pynokio (0x5643). A file is treated as watermarked when at least 15 of the 16 bits are read with clear confidence.

2. A signed provenance record

New files receive a signed, machine-readable provenance record. When a file is downloaded, it comes with the HTTP header X-Pynokio-Synthetic-Media: 1 and a link to its record. The record identifies Pynokio, the type of media, the model or provider, the operation and the time of generation - without revealing the user's prompt or identity.

3. Signed file metadata (coming)

Embedding a signed C2PA manifest directly in downloaded audio files is being introduced. Until it is available, files do not contain it. The full technical description is published in Machine-readable marking.

How to check an audio file

Anyone can check a recording for free, without an account, using the Pynokio Audio Detector:

  1. Open the Audio Detector.
  2. Drop a WAV, MP3, OGG, FLAC, M4A or WebM file (up to 32 MB). The first 60 seconds are analysed.
  3. Read the result. The detector lights up one segment of the wave for every bit of the identifier it reads.

Developers, newsrooms and fact-checkers can call the same free service from software:

curl -F "[email protected]" https://pynokio.com/api/audio-detector/verify

The response contains the result, the identifier bits and the share of the identifier that was read.

How to read the result

ResultWhat it means
Pynokio original confirmedThe file has a valid signed record from Pynokio and its audio has not changed since it was created.
Watermark foundThe Pynokio identifier was read. Strong evidence of Pynokio origin, though not a cryptographic signature.
Possible watermarkPart of the identifier was read, as in a heavily compressed or edited copy. Origin is not confirmed.
No watermark foundThe identifier is not in this audio. This does not prove the file was made elsewhere or by a human.

What a watermark can and cannot prove

The watermark is read reliably after common edits: MP3 at 128 kb/s or more, trimming, background music, reverb, bass or loudness changes. Heavy processing can make it unreadable - telephone-band filtering, loud added noise, aggressive noise reduction, or changes of speed or pitch. Very short clips also carry fewer readable bits.

Remember: a watermark is evidence, not proof of authorship. A missing watermark does not prove a recording is human-made, and a detected one does not say who created it.

Why this matters now

Voice cloning makes fraud such as fake "family emergency" calls and fabricated statements easier. Watermarking gives platforms, journalists and ordinary listeners a way to check. It also reflects the direction of regulation: the EU AI Act requires providers of AI systems that generate synthetic audio, images, video or text to mark their outputs in a machine-readable format so they can be detected as artificially generated. Read more in our guide to voice cloning, consent and the law and in our AI Transparency notice.

What this means for creators

  • Nothing changes in how your audio sounds. The watermark is inaudible.
  • Do not try to remove it. Removing, altering or circumventing the watermark or provenance breaks our Voice Policy.
  • Still disclose. Machine-readable marking does not replace telling your audience when realistic content is AI-generated.

FAQ

Can you hear the watermark?

No. It is designed to be inaudible and does not change how the voice sounds.

Does converting to MP3 remove the watermark?

Not at common quality levels. The watermark is read reliably in MP3 files at 128 kb/s or more.

Is the Audio Detector free?

Yes. It is free for everyone, requires no account, and the uploaded file is deleted as soon as the check is finished.

Can the detector tell if a voice was made with other AI tools?

No. It checks for the Pynokio watermark only. A missing watermark does not mean a file is human-made.

Which Pynokio files carry the watermark?

Generated speech: text to speech, voice clone previews and voice design previews. All new generated files also receive a provenance record.

Got a recording you are unsure about? Check it for the Pynokio watermark - free and without an account.

Open the Audio Detector

Link copied to clipboard!