AI watermark remover GitHub repos: four classes of tool, and the one nobody can write

Repos read, not summarised Answer in the first paragraph Measured on this machine

GitHub's search API reported 1,766 repositories matching "watermark remover" and 414 matching "ai watermark remover" on 2026-09-29. We read the ten with the most stars and ran two of their cleaners. They split into four classes by which layer they touch, and only three of those four are removable. The fourth class is the statistical text watermark — the one Anthropic uses for Claude, the one Google's SynthID-Text builds, the one the Kirchenbauer green-list scheme describes. It is not a byte anywhere in your file, and no tool on GitHub removes it, including this site's. Every repository that addresses that class does the same thing instead: it rewrites your text with a second model. That is a different operation from removing a mark, and the most-starred repository in the category says so in its own README.

The other three classes are real and they come off. A visible overlay drawn into the pixels can be reverse-blended or inpainted. Container metadata — C2PA Content Credentials, EXIF, XMP, ID3 — can be stripped exactly. Invisible Unicode characters can be deleted. What is interesting, and what nobody has written up, is that the tools do not agree on the third one: we diffed four published strip tables and they overlap by as little as 8%, then ran one nine-character probe through all four and got four different answers. All four of them removed the same three characters. Everything after that is a judgement call each project made differently.

FOUR THINGS CALLED A WATERMARK 1,766 GitHub repositories match "watermark remover". They sort by which layer they touch. 1 · Visible overlay, drawn into the pixels removable A corner logo or a sparkle mark. Reverse alpha blend, inpaint, or borrow pixels from another frame. 2 · Container metadata removable C2PA Content Credentials, EXIF, XMP, IPTC, ID3, MP4 boxes. Bytes sitting beside the content, not in it. 3 · Invisible Unicode in text removable, tables disagree Declared code points: 35 (this site), 60, 171, 386. Pairwise overlap 0.08 to 0.53. 4 · Statistical text watermark not removable Claude, SynthID-Text, the green-list scheme. Not a byte in the file. The only repo path is to rewrite the text. One 9-character probe, four tools: 4 different survivor sets, and all four removed the same 3 characters.

What the search actually returns

Searching GitHub for a tool to remove an AI watermark is not a search with a shortage of answers. Queried through the GitHub search API on 2026-09-29, the counts were:

queryrepositories
watermark remover1,766
c2pa1,022
ai watermark remover414
claude watermark150
chatgpt watermark remover15
invisible characters remover8

The shape of that table is the first useful thing it tells you. The query with the most results is the vaguest one, and the queries that name a specific provenance scheme return almost nothing. chatgpt watermark remover returns 15 repositories. invisible characters remover returns 8. Whatever people are looking for when they type those phrases, the supply is thin, and thin supply in a search result is a sign that the thing being asked for is hard to build rather than that nobody has thought of it.

Star rank does not sort by topic

The ten most-starred repositories matching those queries, as the API ranked them on 2026-09-29:

repositorystarswhat it is
guillaumemeyer/watermarks-remover23,041A text and file cleaner that also runs an agent rewrite pass
YaoFANGUK/video-subtitle-remover13,090Removes hard-coded subtitles and burned-in text from video
wiltodelta/remove-ai-watermarks5,708Pixel regeneration profiles for image and video marks
GargantuaX/gemini-watermark-remover5,612Reverse alpha blending on Gemini's visible sparkle mark
D-Ogi/WatermarkRemover-AI1,993Florence-2 plus LaMA inpainting over a detected watermark
tomateo1reg/sora2-watermark-remover-web-gui833A web interface aimed at Sora output
lxulxu/WatermarkRemover820Batch removal of a fixed-position mark from video
dinoBOLT/Gemini-Watermark-Remover371A browser extension for the same Gemini mark
whitelok/watermark-remover310Classical algorithm for a fixed-position mark
allenk/VeoWatermarkRemover277Reverse alpha blending for a Veo video mark

Two things stand out. First, the second-most-starred repository in the set is not an AI watermark tool at all — it removes burned-in subtitles, which is a different problem that happens to share the word "watermark". Second, only one entry in the top ten is about text. Everything else is about pixels. If you arrived at GitHub looking for something to do with the text a language model produced, almost all of the stars are pointing somewhere else.

Stars measure how many people wanted the thing, not whether the thing does what its name says. That is not a criticism of any of these projects — it is a warning about how to read a search result, and it is why the rest of this page is about reading what each project says about itself.

The four classes, and the one nobody can write

Sort the repositories by what they touch rather than by what they are called and the field collapses into four classes. Three of them are file operations. One of them is not.

classwhere the mark livesverdict
Visible overlayPixels, on top of the image or videoRemovable. Blend, inpaint, or borrow pixels from a frame where the mark is not covering that spot.
Container metadataStructures beside the content: C2PA boxes, EXIF, XMP, ID3Removable exactly. Nothing is re-encoded.
Invisible UnicodeCharacters in a text stringRemovable. Which characters count as invisible is a judgement call — see below.
Statistical text watermarkNowhere in the file. It is a bias in how tokens were chosen.Not removable by any tool. Detection needs the issuer's key.

Anthropic describes its own text watermark in exactly those terms. The company's announcement says the scheme is a statistical pattern in token selection and that "Nothing is added to the text and there are no hidden characters" — which is a sentence worth reading twice, because it is the issuer telling you there is no byte to delete. Detection requires Anthropic's key. The same architecture appears under other names elsewhere: Google's SynthID-Text uses a tournament between candidate tokens, and the academic green-list scheme biases which tokens are eligible. All three are keyed, all three are statistical, and none of them writes anything into your file. Why a character cleaner cannot reach that layer →

This is the part the category keeps fudging, so it is worth being blunt: a repository that removes a C2PA box from a JPEG and a repository that claims to remove an AI watermark from text are not doing the same kind of work, even when both are named "ai-watermark-remover". The first is deleting bytes. The second has no bytes to delete.

What the repositories say about the class they cannot touch

You do not have to take our word for any of that, because the projects say it themselves. These are quotes from the repositories' own documentation, read on 2026-09-29. Short excerpts, attributed.

guillaumemeyer/watermarks-remover (23,041 stars) publishes a two-layer table in its README. Layer A is the character and metadata pass. Layer B is listed against the row that reads Statistical (token-sampling) text watermarks, and the method named for it is Agent rewrite plus an optional rewrite_text.py hook. The same README is explicit about the limit of that hook:

"What a hook cannot do. No hook can rewrite the assistant's chat message"

Its detector table also lists the research harness it can drive, and marks it as "same-config-only — not a vendor oracle" — a project being careful not to claim it can verify a mark it does not hold the key for.

yasir-mo/AI-watermark-remover-GUI (84 stars) draws the same split in ASCII, and its diagram labels the statistical row "rewrite, probabilistic". The README then explains the mechanism without dressing it up:

"So Layer B rewrites the text with a second model. That model samples from its own distribution, using its own key or none, so the original pattern does not [survive]"

It follows that sentence with the cost, in four words: "Rewriting costs quality." That is the honest version of what every "removes the statistical watermark" claim actually delivers — a new text, produced by a different model, which is not the text you had.

mikiane/claude-watermark-cleaner (147 stars) is a single Python file whose docstring states the limit before you run it: cleaning is limited to Unicode and cannot remove a watermark based on word choice. We ran it. In --engine clean mode it printed, unprompted, that Unicode cleaning alone does not remove a statistical watermark based on word choice. Its whole design is "clean, then rephrase" — two steps, because the first one cannot do the job alone.

Neeeophytee/ai-watermarks-reality-check (14 stars) is the most useful repository in the set and the least starred, because its entire purpose is to report what cannot be checked. Its detector table lists synthid-text as never available in the build — not configured without keys, unsupported with them — and its README states the reason plainly:

"Keyed model-level watermarks bias token sampling with a secret key; detection requires that key. A third party cannot check them, and this pack says so rather than guessing."

It also declares a deliberate non-goal: statistical and stylometric AI-text classifiers are not implemented and will not be added, because their false-positive rates make them unsafe for accusing a person of AI authorship.

GargantuaX/gemini-watermark-remover (5,612 stars) has a Limitations section that says, in one line, what the name does not promise: it "does not remove invisible/steganographic watermarks". A 5,600-star project naming its own boundary is a good sign. mertizci/noai-watermark (193 stars), which does attempt to disrupt pixel-level marks by regenerating the image, is equally careful: it notes that a successful regeneration run "does not establish the absence of an invisible watermark", and that whether a mark remains requires an independent compatible detector which the tool does not provide. wiltodelta/remove-ai-watermarks (5,708 stars) says the same thing in its own way — there is no local SynthID pixel detector in the package, so it reports a regeneration rather than a per-file verdict.

Put those side by side and a pattern shows up that the search results hide: the projects that are closest to the statistical layer are the ones most careful about claiming it. The vagueness lives in titles and descriptions, not in the code.

We diffed the character tables

Class three — invisible Unicode — is where the field is genuinely open, because "invisible" is not a property Unicode defines. It is a list each project writes by hand. So we read the lists out of the source files and compared them. Every set below was copied from the named file in the named repository on 2026-09-29; nothing is inferred.

projectwhere the list livescode points
this sitecleaner.js, ZERO_WIDTH + BIDI + CONTROL35
guillaumemeyer/watermarks-removerservice/scripts/text_unicode.py, STRIP_CODEPOINTS60
mikiane/claude-watermark-cleanerclean_claude_watermark.py, SUSPICIOUS_RANGES171
yasir-mo/AI-watermark-remover-GUIscrubai/unicode_tables.py, nine named groups386, plus the entire private use area

Pairwise overlap, as a Jaccard ratio of the declared sets with the private use areas excluded:

pairoverlap
this site × guillaumemeyer/watermarks-remover0.53
guillaumemeyer × mikiane0.23
mikiane × yasir-mo0.30
guillaumemeyer × yasir-mo0.12
this site × mikiane0.19
this site × yasir-mo0.08

The highest agreement in the whole set is 0.53. Two projects that both claim to strip invisible characters from AI text agree on about half of what that means. The lowest is 0.08. This is not a subtle difference of implementation — it is four different definitions of the word "invisible", shipped under one shared name.

Two divergences are worth naming, because they show the disagreement is a design choice rather than an oversight. U+2028 and U+2029, the line and paragraph separators, are declared by this site and by none of the three. In the other direction, all three declare the variation selectors U+FE00–U+FE0F and the tag characters U+E0020–U+E007F, and this site does not.

A bigger list is not automatically the better tool, and one of the projects says so in its own flag help. yasir-mo's cleaner carries an option documented as "do not preserve load-bearing joiners (breaks emoji and Persian/Indic text)" — which is an admission that its thoroughness and its correctness pull in opposite directions. The same file marks its homoglyph folding as aggressive-only, with the comment that those are visible characters and stripping them changes text that may legitimately be Cyrillic or fullwidth. The variation selectors it strips in its default mode are the ones that make ❤️ render as an emoji rather than a black heart. Removing them is a real edit to the text, made silently, in the name of watermark removal.

One probe string, four different answers

Tables are abstract, so we put nine different invisible characters into one sentence and ran it through each cleaner. The sentence was Ship the quarterly report today before noon to the client. with one invisible character planted after each of the first nine words. Two of the four cleaners were executed — mikiane's script directly, and this site's cleaner.js through Node. The other two packages are not installed on this machine, so their published tables were applied as declared, and this page says which is which rather than blurring the two.

plantedmikiane
(ran it)
this site
(ran it)
guillaumemeyer
(table)
yasir-mo
(table)
U+200B zero width spaceremovedremovedremovedremoved
U+00AD soft hyphenremovedremovedremovedremoved
U+200D zero width joinerremovedremovedremovedremoved
U+2064 invisible plusremovedremovedremovedremoved
U+180E Mongolian vowel separatorremovedremovedremovedkept
U+3164 Hangul fillerremovedkeptremovedkept
U+FE0F variation selector-16keptkeptremovedremoved
U+2028 line separatorkeptremovedkeptkept
U+E0041 tag Latin capital Aremovedkeptkeptremoved
survivors2323

All four removed the same three characters. Everything past that is disagreement. One character — U+2064 invisible plus — was removed by all four, but only after we fixed this site: before this measurement, our own table was missing it, and the probe caught it. That is the whole argument for running a comparison instead of reading one.

The consequence is practical. If you paste a passage into one of these tools, get a clean report, and paste it into another, the second one may find something the first one left behind. A tool reporting "0 invisible characters found" is reporting against its own list, not against Unicode.

What this site's table covers, and what it deliberately does not

Our list is the shortest of the four at 35 code points, and three of the characters the probe found surviving were left out on purpose rather than by accident. It is fair to hold that against us, so here is the reasoning, and here is what we changed today.

We would rather ship 35 code points we can defend than 386 that quietly rewrite emoji and Mongolian. If your text is English prose and you want maximum aggression, a wider table is the right call and two of the projects above offer one — just know that "wider" is a trade, not a free upgrade. Run this cleaner locally, with the same 35 code points →

How to tell which class a repo is in, in about a minute

You do not need to read the code. Open the README and search it for four words, in this order.

what you findwhat it means
rewrite, paraphrase, humanize, or an Ollama / API-key setup stepIt is handling the statistical layer by replacing your text. It does not remove a mark; it produces a different document. Expect the wording to change.
C2PA, EXIF, metadata, ID3Container class. Real, exact, and verifiable — check that it says it does not re-encode.
alpha, blend, inpaint, delogo, regeneratePixel class. Works on visible marks. Ask what happens when the mark sits over moving content.
zero-width, invisible, Unicode, confusableCharacter class. Then check whether the README warns about emoji or right-to-left languages, because that is where a wide table does damage.
none of the four, and no limitations sectionThe most common case, and the one to skip.

The single strongest signal in the whole survey was a Limitations section. Every project that had one was honest about the statistical layer, and the projects with 5,000 stars and up had the most specific ones. A repository that names what it cannot do has usually thought about the question; a repository that claims everything has usually renamed it.

And whatever you pick, the thing it will not do is the thing the name promises. The statistical text watermark is not in your file, so no repository can take it out. If a tool offers to, it is rewriting your prose, and the honest ones put that on the label. What this site's tool actually removes →

Frequently asked questions

Is there an AI watermark remover on GitHub that actually works?

It depends which of the four classes you mean, and that is not a dodge. For visible overlays and for container metadata there are good repositories and they do what they say. For invisible Unicode there are at least four and they disagree with each other, so the answer depends on which characters you are worried about. For the statistical text watermark — Claude's, SynthID-Text's, the green-list scheme's — there is none, and there cannot be one, because there is no byte in the file to delete and detection needs the issuer's key. Any repository claiming that last one is rewriting your text with a second model, which is a content change rather than a removal.

Why does the most-starred AI watermark remover rewrite my text instead of cleaning it?

Because rewriting is the only operation available at that layer. If a watermark is a statistical bias in which tokens were chosen, the way to disturb it is to choose tokens differently — which means generating new text with a model that samples from its own distribution. That is why the biggest repository in the category lists Agent rewrite against its statistical row, and why another one labels the same row rewrite, probabilistic and adds that rewriting costs quality. The wording you get back is not the wording you gave it.

Do all these tools strip the same invisible characters?

No. We read four published strip tables and compared them: this site declares 35 code points, and the other three declare 60, 171 and 386 plus the entire private use area. Their pairwise overlap runs from 0.53 down to 0.08. Then we planted nine invisible characters in one sentence and ran all four: the survivors were 2, 3, 2 and 3, and the only characters all four removed were the same three. A "0 invisible characters found" report is a statement about that tool's list, not about Unicode.

Is a GitHub repo safer than an online watermark remover site?

Not automatically, but it is more checkable, and that is the real advantage. A repository lets you read the strip table, see whether it re-encodes, and confirm whether it makes a network call. An online tool asks you to trust a claim. If you want the strongest version of that argument, look for a tool that runs entirely client-side or as a local CLI, so the file never leaves the machine — this site's own cleaner runs in the browser and as a local MCP server for that reason. How the local path works →

Can a GitHub tool remove a SynthID watermark?

Not by removing anything. SynthID-Text biases token selection with a key, and verification needs the same key, so a third party cannot even check for it, let alone delete it — one of the repositories we read states exactly that and refuses to guess. What repositories do instead is regenerate or rewrite the content so the original statistical pattern no longer describes it. That can change the content substantially, and the careful projects say that a successful run does not prove the mark is gone, because they do not have a detector to prove it with.

Which of these repositories should I use for photos?

If the mark is visible in the picture, the pixel tools are the right family: they reverse-blend or inpaint over a known region, and several of them are mature. If what you want gone is the metadata — C2PA Content Credentials, GPS, camera serial numbers — you do not need any of them, because that is a container operation and one pass removes it without touching a pixel. Our own image path does that in the browser and lists exactly which segments it dropped and which it kept. What each metadata method leaves behind →

Does this site's cleaner do anything the GitHub projects do not?

It handles image metadata and video containers, which most of the text-focused repositories do not, and it runs entirely on your device with no upload and no account. On the character layer specifically it is deliberately narrower than the others, for the reasons set out above. It cannot touch the statistical watermark, in text or anywhere else, and no tool can — the honest position is that the layer people most want removed is the one layer that is not a file operation at all.