AIwatermark.studio

Invisible characters from ChatGPT and Claude: find them and strip them

Updated 2026-06-03

Text copied from an AI often carries characters that do not render: zero-width spaces, joiners, variation selectors, tag characters. They survive copy-paste, travel from a word processor into a CMS, and are enough to identify where a passage came from. The good news: these are the easiest marks to remove, and removal is verifiable.

01

Which characters are involved

  • U+200B zero-width space, U+200C and U+200D joiners: invisible between two letters, but present in the bytes.
  • U+FEFF byte order mark, used as an invisible space mid-text.
  • U+2060 to U+2064: invisible operators with no rendering at all.
  • Variation selectors U+FE00–U+FE0F and U+E0100 onwards: designed to pick a glyph form, repurposed as a data channel.
  • Tag characters U+E0000–U+E007F: the stealthiest channel, able to encode entire strings completely invisibly.
  • Bidirectional controls (U+202A–U+202E, U+2066–U+2069) and space homoglyphs (thin space, narrow no-break space) standing in for a normal space.
02

How to tell whether your text contains them

By eye, you cannot — that is the whole point. Three practical hints: your editor's character count is higher than what you can see; a find-and-replace on a word fails even though the word is right there; the text wraps in a strange place.

The reliable method is analysis: a tool that walks the code points and flags every Cf-category character and every variation selector. A good report gives you the position, the code point and the class of each occurrence, rather than a vague "suspicious text".

03

The trap of brute-force cleaning

Deleting every invisible character breaks perfectly legitimate text. U+200D is required after a composed emoji: without it a family becomes four separate figures and a regional flag turns into two letters. U+200C separates letters in Persian, Arabic and Devanagari: removing it changes the spelling.

The right rule is contextual. Keep load-bearing invisibles: a joiner after an emoji base, a non-joiner inside a word in a script that needs it, the tag characters of a valid flag. Strip everything else. A "paranoid" mode can remove all of it, but it must be an explicit choice, never the default.

04

Clean, then verify

After cleaning, paste the clean text back into the analyser: if the report is empty, no carrier is left. That round-trip check is what separates a real removal from a marketing claim.

One important caveat: removing invisible characters does not make text undetectable as AI writing. That is a different, statistical family of marks living in the word choices themselves. The two problems are handled separately.

Frequently asked questions

Does ChatGPT really add invisible characters?

Text from AI assistants regularly contains zero-width spaces and invisible sequences, whether they come from the model, the web interface rendering or a copy intermediary. Whatever the source, they are detectable and verifiably removable.

Will cleaning break my emoji?

No, provided the tool preserves load-bearing invisibles: the ZWJ after an emoji base, flag tag characters and the orthographic marks of scripts that depend on them. Only an explicit paranoid mode removes those.

Do Word or a grammar checker strip these characters?

Not reliably. Most word processors keep them as-is through copy-paste, and converting to plain text does not remove them either, since they are part of the text.

Clean your file or your text

Detection and cleaning inside your browser: no upload, no account, no quota. Your files never leave your device.

Open the studio