An audio file can carry two unrelated things, and only one of them is metadata. Tags — ID3v2 frames in an MP3, a moov/udta tree in an M4A, Vorbis comments in a FLAC, a RIFF INFO list in a WAV, and any C2PA Content Credential — sit outside the samples and come off. A signal mixed into the samples does not. We measured both on the same file. A tone sitting 50 dB below the music read −50.0 dBFS before anything was done to it, −50.0 dBFS after a tag strip, −50.0 dBFS after this site's own cleaner, and −50.4 dBFS after an MP3 re-encode at 64 kbps. The tags went from 12 regions to 0 in the same pass. The tone did not move.
The second number is the one people get wrong, so it is worth being exact about it. Re-encoding is not a removal method; it is a filter, and a filter only reaches what is inside its band. Drop the same file to 56 kbps and the tone reads −87 dBFS — identical to the control file that never had a tone in it — because at that bitrate the codec's bandwidth stops below 15,500 Hz. Put the same tone at 3,000 Hz instead and it survives all the way down to 32 kbps. Nothing cleaned anything in either case.
Search for an audio watermark remover and every result describes the same thing: you upload a file, the site strips something out, you download it. What none of them say is which layer they touched, and that is the only question that matters. A tag strip is a container operation — it rewrites the region of the file that holds descriptive data and leaves the audio bitstream alone. It cannot see a mark that lives in the audio, because the audio is not where it is looking.
The two layers also fail in different ways. Tags are easy to remove and easy to re-add, and they travel with the file when you copy it. A signal in the samples is hard to remove and hard to detect, and it travels with the audio through a copy, a format conversion, and — depending on where it sits — through a lossy re-encode. If you do not know which one you are dealing with, "remove the watermark" is not a task you can finish or verify.
Before asking what comes off, it is worth seeing what is in there. We built four files from one 3-second tone and tagged each the way a recorder or a music service would, then read the containers directly with a standard-library script rather than trusting a tool's success message.
| File | Size | Metadata regions | What they are |
|---|---|---|---|
plain.mp3 | 73,658 B | 10 | One ID3v2.3 tag, 308 bytes, holding 9 frames: title, artist, album, year, genre, a custom field, copyright, "encoded by", and the encoder's own settings string |
plain.m4a | 52,288 B | 11 | A moov/udta tree → meta → ilst carrying 8 iTunes-style tags: title, artist, album, date, encoder, comment, genre, copyright |
plain.flac | 44,058 B | 2 | A Vorbis comment block (299 bytes) and an 8,192-byte padding block |
plain-tagged.wav | 264,914 B | 1 | A RIFF LIST/INFO chunk, 262 bytes |
plain-c2pa.m4a | 52,380 B | 12 | The same M4A plus a top-level uuid box, 92 bytes — the position a C2PA provenance manifest occupies in an ISO base media file |
Two things are worth pulling out of that table. First, the formats disagree about where tags live: an MP3 keeps them in one block before the audio, an M4A keeps them in a tree inside moov after the audio, a FLAC keeps them in a metadata block with 8 KB of padding behind it, and a WAV keeps them in a RIFF chunk that most people never look at. A tool that handles one is not automatically handling the others.
Second, C2PA is designed to sit in exactly these places. The specification's list of embeddable formats includes ID3 and RIFF alongside ISO BMFF — so a Content Credential on an audio file is a container object, in the same family as the tags next to it. That matters below: it is the part of an AI audio provenance chain that a cleaner can actually reach.
We then took the C2PA-bearing M4A and the tagged MP3 and ran them through the three things people actually do, printing the byte size, the regions left behind, and the MD5 of the decoded samples after each one. The sample hash is the control: it says whether a method touched the audio or only the container.
| Method | File | Bytes | Metadata regions left | Sample MD5 | Audio re-encoded? |
|---|---|---|---|---|---|
| — baseline | plain-c2pa.m4a | 52,380 | 12 | 68b01ab328acbfdc | — |
A. remux, tags droppedffmpeg -map_metadata -1 -c copy | A-remux.m4a | 51,979 | 4 | 68b01ab328acbfdc | no |
| A. same, on the MP3 | A-remux.mp3 | 73,394 | 2 | c895b8a143d88ee7 | no |
B. this site's cleanervideo-cleaner.js | B-site.m4a | 52,380 | 0 | 68b01ab328acbfdc | no |
C. full re-encodeffmpeg -c:a libmp3lame -b:a 192k | C-reencode.mp3 | 73,658 | 10 | 2a49095178df120c | yes |
Read that table the way the hashes tell you to read it, and three results fall out that the usual advice gets backwards.
-map_metadata -1 took the regions from 12 to 4, not to zero: a moov/udta/meta/ilst skeleton survives carrying one tag, ©too = Lavf63.1.101. On the MP3 the tag was rewritten from ID3v2.3 to ID3v2.4 and a TSSE frame was left in place. Both files now name the tool that cleaned them. That is not a watermark, but it is metadata you did not write, sitting in the file you are about to send.video-cleaner.js reported 9 metadata regions before and 0 after, hasC2PA true then false, and its own six checks all passed — including "mdat payload identical (50649 bytes)" and "output length unchanged (52380 bytes)". It rewrites each metadata region in place as a zeroed free box, so nothing shifts and the audio payload is byte-for-byte the same. Same size, different contents.TSSE frame, and the sample MD5 moved from c895b8a143d88ee7 to 2a49095178df120c. If your goal is privacy rather than sound, that is the worst of both columns: you degraded the audio and kept the metadata. Converting the format behaves the same way — an MP3 converted to WAV arrived with its 262-byte RIFF INFO chunk intact, and an M4A converted to MP3 arrived with every tag plus three new custom fields.None of the above touched a signal inside the audio, and that is measurable rather than theoretical. We took the same tone and mixed a second tone into it, 50 dB down, then ran every method again and measured the level of that second tone with a one-frequency detector built on the Goertzel algorithm. The control is the file with no second tone in it, which is how you can tell the detector is not inventing a reading.
| What was done | Mark at 15,500 Hz | Mark at 3,000 Hz | The 440 Hz audio |
|---|---|---|---|
| control: no mark in the file | −86.8 dBFS | −88.3 dBFS | −20.0 dBFS |
| mark present, nothing done | −50.0 | −50.1 | −20.0 |
after ffmpeg tag strip (-c copy) | −50.0 | −50.1 | −20.0 |
| after ffmpeg tag strip on the MP3 | −50.3 | −50.4 | −20.3 |
| after this site's cleaner (M4A) | −50.0 | −50.1 | −20.0 |
| after re-encode to 192 kbps | −50.3 | −50.4 | −20.3 |
| after re-encode to 128 kbps | −50.4 | −50.5 | −20.4 |
| after re-encode to 96 kbps | −50.4 | −50.5 | −20.4 |
| after re-encode to 64 kbps | −50.4 | −50.5 | −20.4 |
| after re-encode to 48 kbps | −86.6 | −50.5 | −20.4 |
| after re-encode to 32 kbps | −87.0 | −50.5 | −20.4 |
The 15,500 Hz column holds steady through every tag operation and every re-encode down to 64 kbps, then collapses at 48 and 32 kbps to −86.6 and −87.0 dBFS — the control's own noise floor. The 3,000 Hz column never collapses at all. The 440 Hz column, which is the audio itself, stays at roughly −20 dBFS throughout, so nothing went wrong with the files.
To pin down what actually happened, we encoded the marked file at six bitrates and probed the level at several frequencies in each:
| MP3 bitrate (mono) | 15,000 Hz | 15,500 Hz | 16,000 Hz |
|---|---|---|---|
| 40 kbps | −101 dBFS | −87 dBFS | −111 dBFS |
| 48 kbps | −100 | −87 | −109 |
| 56 kbps | −104 | −87 | −114 |
| 64 kbps | −91 | −50 | −107 |
| 80 kbps | −92 | −50 | −107 |
| 96 kbps | −91 | −50 | −112 |
The boundary sits between 56 and 64 kbps for a tone at 15,500 Hz in mono. Above it the tone is there at full strength; below it the codec's bandwidth no longer reaches that frequency at all, and no amount of "cleaning" is involved. The rule is about where the signal sits, not about re-encoding as a technique. A mark placed inside the band the codec keeps is untouched by the re-encode; a mark placed above the band is destroyed along with everything else up there. Neither outcome is a removal method you can point at.
The AI audio services do both things, and their own documentation says so.
Google's SynthID page describes the audio layer as something embedded in the audio itself: “SynthID embeds a watermark into any audio generated or published through our AI music generation model Lyria or the podcast generation feature of Notebook LM. It's inaudible to the human ear, and can't be altered by common modifications like adding noise, MP3 compression, or changing the speed of the track.” That is a claim about the samples layer, and it is consistent with what we measured here: MP3 compression does not touch a signal that sits inside the codec's band. The same page states the general form — “SynthID embeds digital watermarks directly into AI-generated images, audio, text or video.”
OpenAI's provenance announcement describes both layers as complementary, and its July 31, 2026 update extended the audio side: “We're expanding this work beyond images. Supported audio generated with OpenAI tools, including through ChatGPT and the OpenAI API, now includes SynthID watermarking. Our public verification tool will now allow for verification of supported audio files in addition to images.” The same announcement is blunt about the metadata layer: “metadata is not foolproof. It can be stripped, lost through uploads and downloads, or broken by transformations like file format changes, resizing, or screenshots.” It also records the earlier, narrower use of the word — “we have used visible watermarks in Sora and an audio watermark in Voice Engine” — which is a third thing again: a signal added deliberately to a specific product's output, not a general property of all audio.
So the honest summary for an AI-generated track is this. The Content Credential attached to it is a container object in an ID3 tag, a RIFF chunk or an ISO BMFF box, and it comes off the way every other tag on this page came off. The embedded signal is in the audio, and a tag tool cannot reach it; whether a re-encode reaches it depends entirely on the band it occupies, which is a property of the service's implementation and not something this page can generalise about. And a third layer has nothing to do with either: detection is not binary. OpenAI's own wording is that “no detection method is foolproof”, and that when its tool finds nothing it “will not make a definitive conclusion”. A file that reads clean is not evidence of anything.
This is the same shape as the text case, and it is worth naming the parallel because the two get conflated constantly. Anthropic's text watermark is statistical — its own announcement says “nothing is added to the text and there are no hidden characters” — so no character cleaner can remove it, and this site does not claim to. What a character cleaner removes from text is a different layer: invisible and zero-width characters, directional marks, hidden attributes. Audio has the identical split, just with different names for the parts. In both cases the removable layer is the container and the unremovable layer is the content. The text version of that distinction →
For the container, match the method to the format rather than reaching for the nearest online tool:
ffmpeg -map_metadata -1 -c copy is the right tool, with the caveat measured above — it leaves its own encoder string in the file. If that matters to you, run it and then check.For the signal, the honest answer is that there is no general method, and any page that offers one is either talking about tags or guessing. What you can do is find out what you are dealing with before deciding: measure the file, and treat a null result as a null result rather than as a clean bill of health. The video container works the same way →
Neither measurement needs anything installed beyond ffmpeg and Python. This is the container report — it reads the bytes and names the regions, for MP3, M4A/MP4, FLAC and WAV:
python audio_meta_report.py track.m4a track.m4a (52288 bytes) moov/udta (user data) offset 51882 406 bytes moov/udta/meta (iTunes-style) offset 51890 398 bytes moov/udta/meta/ilst offset 51935 353 bytes moov/udta/meta/ilst/©nam (title) offset 51943 47 bytes moov/udta/meta/ilst/©ART (artist) offset 51990 38 bytes ... 11 metadata region(s) across 1 file(s)
And this is the signal check. It reports the level of one frequency, in dBFS, using the Goertzel algorithm — one pass, no FFT library, and a window that is a whole number of cycles of the target frequency so nothing is smeared:
python audio_tone_check.py 15500 track.wav track-clean.wav target frequency 15500 Hz, window = a whole number of its cycles file level dBFS vs control track.wav -50.0 track-clean.wav -50.0 +0.0 dB
Run it on a file with no tone at that frequency and you get the noise floor instead — about −87 dBFS in the fixtures used here. That difference, roughly 37 dB, is the whole test: it tells you whether the frequency you named is present, without you having to trust a tool's summary of what it did.
Split the question in two and it has two different answers, which is why the general version is unanswerable. The Content Credential attached to an AI-generated track is container metadata — it rides in an ID3 tag, a RIFF chunk or an ISO BMFF box — and that comes off, as measured above. An inaudible signal embedded in the samples does not come off, and no tag tool will ever reach it, because the tool is not looking at the audio. If a site promises to remove an "AI watermark" from audio without saying which layer it means, it is describing the first one and letting you assume the second.
No. We converted a tagged MP3 to WAV and it arrived carrying its 262-byte RIFF INFO chunk; a tagged M4A converted to MP3 arrived with every tag intact plus three new custom fields the converter added. A format conversion is a re-encode, and re-encoding copies metadata rather than dropping it. If you want the tags gone, the strip has to be a deliberate step.
No, and they are not even in the same part of the file. Metadata is descriptive data stored in a container structure alongside the audio: an ID3v2 frame in an MP3, an ilst entry inside moov/udta in an M4A, a Vorbis comment in a FLAC, a RIFF INFO field in a WAV. A watermark of the embedded kind is a signal inside the audio samples themselves. One is a label attached to the file; the other is part of the sound. A tag editor shows you the first and cannot show you the second, which is why people who check a file's tags, see nothing, and conclude there is no watermark are looking in the wrong layer.
Not in the sense images do. EXIF is an image and camera format, and it is why a photo can carry GPS coordinates, a camera serial number and a thumbnail. Audio has its own equivalents: ID3 for MP3, Vorbis comments for FLAC and Ogg, RIFF INFO for WAV, and the iTunes-style ilst tree for M4A and MP4. The fields overlap in spirit — who made it, with what, when — but the containers and the parsers are different, so a photo tool will not read an MP3 and vice versa. What each photo method leaves behind →
For an M4A, drag it into the tool on this site: it reads the box tree, lists every region it found, frees them in place, and hands the file back without re-encoding. For an MP3, FLAC or WAV, one command does it without touching the audio — ffmpeg -map_metadata -1 -c copy in.mp3 out.mp3 — and the thing to remember is that the output will still carry an encoder string written by ffmpeg itself. That is one tag, not zero, and the only way to know is to read the result rather than assume it.
On M4A and MP4-family audio, yes, and we measured it rather than assuming it: 9 metadata regions became 0, hasC2PA went true to false, the audio payload was byte-for-byte identical, and the file size did not change because the regions are freed in place. On MP3, FLAC and WAV it does not, and we would rather say so than let you upload a file and watch nothing happen. It also cannot touch a signal embedded in the samples in any format, for the reason this whole page is about. What the tool does and does not remove →
Name a frequency and measure it. The audio_tone_check.py command above decodes the file to raw samples with ffmpeg and then runs a Goertzel filter at exactly the frequency you asked for, printing the level in dBFS. A reading far above the file's noise floor means something is there. The limit is that this is a targeted test, not a scan: it answers "is there a tone at 15,500 Hz", not "is there a watermark somewhere". A spread-spectrum mark would not show up as a single peak at all, which is the point of spreading it. Treat a negative result as "nothing at this frequency", never as "nothing in the file".