Invisible letter: four of them are real letters, and every letters-only rule waves them through

Short answer: of the twelve code points Unicode names as fillers, exactly four are letters — U+3164 HANGUL FILLER, U+1160 HANGUL JUNGSEONG FILLER, U+115F HANGUL CHOSEONG FILLER and U+FFA0 HALFWIDTH HANGUL FILLER. They print nothing and they are not whitespace, which is the combination that makes them useful: \p{L} accepts them, isalpha() accepts them, and no trim will delete them. The zero-width space you were handed instead is none of those things — it is a format character, it fails every letters-only test, and it measures 0.00 px where a filler reserves a whole character cell.

Not a letter, for contrast:

One character, three bytes as UTF-8, no ink on screen. U+3164, U+1160 and U+115F each advance the cursor by 16 px in a 16 px font; U+FFA0 advances it by 8 px, so it is the one to use where a blank name is pushing a layout around. All four pass ^\p{L}+$. None of them is removed by trim().

Everything below was measured, not looked up: Chromium via Playwright for the rendering and form columns, Python 3.13.14 (Unicode 15.1.0) and Node 22.22.2 (Unicode 17.0) for the category and regex columns.

The four, measured

Twelve Unicode characters carry the word FILLER in their name. Four are letters, seven are punctuation and one is a combining mark. Only the four letters behave like a character you typed: they occupy a cell, they survive a trim, and they answer true to every “is this a letter” question anyone thinks to ask.

Code pointNameCat.Width / charInk paintedBytesNFKCJS variable name
U+3164HANGUL FILLERLo16.00 px0 px3U+1160legal
U+1160HANGUL JUNGSEONG FILLERLo16.00 px0 px3unchangedlegal
U+115FHANGUL CHOSEONG FILLERLo16.00 px0 px3unchangedlegal
U+FFA0HALFWIDTH HANGUL FILLERLo8.00 px0 px3U+1160legal
U+200BZERO WIDTH SPACECf0.00 px0 px3unchangedsyntax error
U+E0054TAG LATIN CAPITAL LETTER TCf0.00 px0 px4unchangedsyntax error
U+FFFCOBJECT REPLACEMENT CHARACTERSo16.00 px1070 px3unchangedsyntax error

The ink column is the one nobody publishes. It is a pixel count of the rendered glyph against a blank region of the same size: ten copies of U+3164 in a 16 px font measured 160 px wide and painted zero non-white pixels. The character reserves a full cell and puts nothing in it. That is the whole trick, and it is also the tell — anything with a background, a border or an underline behind a blank name will show a 16 px gap where the eye expects none.

U+FFFC is in the table because it is the usual next guess and it fails. It is the character meant to stand in for a missing object, it is printable, and in Chromium it paints a visible box: 1070 non-white pixels for ten copies, against 0 for the fillers and 492 for a Cyrillic а. Counting painted pixels is the only test that separates the four real invisible letters from an ordinary glyph your font happens to lack.

The letters-only rule is the hole, not the fix

The advice given whenever somebody asks how to accept names in any alphabet is to match \p{L}. It is in the accepted answers, it is the obvious reading of “any letter”, and it lets all four of these in. Measured in Chromium with a real form, a field carrying pattern="\p{L}+" validates a value of U+3164, U+1160, U+115F or U+FFA0, and rejects U+200B, U+00A0, U+FEFF and the tag characters. A field carrying pattern="[A-Za-z]+" rejects everything invisible — and every name that is not written in the Latin alphabet with it.

That is the fork, and there is no clever third option in the middle of it. Either your rule accepts exactly the characters somebody can type on a US keyboard, or it admits every character Unicode files under Letter, including four that nobody can see. The rules that look like a compromise are not: [\p{L}\p{M}] accepts the four as well, because the marks were never the problem.

One rule, two languages, opposite answers

Write the same check in Python and in JavaScript and you get different verdicts on the same string, because the two languages do not mean the same thing by the same pattern.

RuleU+3164U+1160U+115FU+FFA0U+200BU+E0054
Python re.fullmatch(r"\w+")passpasspasspassfailfail
JavaScript /^\w+$/failfailfailfailfailfail
Python s.isalpha()passpasspasspassfailfail
JavaScript /^\p{L}+$/upasspasspasspassfailfail
Python s.isprintable()passpasspasspassfailfail
Python s.isidentifier()passpasspasspassfailfail
either language [A-Za-z0-9_]+failfailfailfailfailfail

Python’s \w is Unicode-aware and matches any letter, so a blank name is a word. JavaScript’s \w is [A-Za-z0-9_] and nothing else, so the same blank name is not. Both are documented, both are correct, and a stack that validates in the browser with \w and again on the server with \w will reject the name in one place and accept it in the other.

The identifier row is the sharpest version of the same thing. All four fillers are legal JavaScript variable names on their own — var followed by U+3164 parses. It has been called a bug in the identifier rules and it has been defended as the price of letting people write code in their own script. Either way it is the property being relied on here: the character is a letter to the language, not to the reader.

The characters named like letters are not letters

The tag block, U+E0020 to U+E007F, is where the naming goes the other way. Its characters are called TAG LATIN CAPITAL LETTER T, TAG LATIN SMALL LETTER A and so on down the ASCII range, and they are not letters. Their category is Cf. \p{L} does not match them. They cannot start or continue a variable name. They cost four bytes and two UTF-16 code units each, against three bytes and one unit for a filler, and they do something no filler does: they attach to the character in front of them.

Measured with the grapheme segmenter in Chromium, the string a + U+E0054 + b breaks into two clusters, not three — the tag character is part of the a. The same test on U+3164 gives three. So a single backspace after a visible letter with a tag character behind it takes the letter with it, and a counter that reports “characters” by counting clusters reports one fewer than a counter that counts code points. The fillers are ordinary standalone characters; the tag characters are passengers.

That difference also decides what your cleanup finds. The popular one-liner that strips [\u200B-\u200D\uFEFF] catches the zero-width space and the byte-order mark and misses all four invisible letters, because letters are not in its range. \p{C} and \p{Cf} catch the tag characters and the zero-width family and miss all four invisible letters, for the same reason in reverse. Only a sweep for anything outside printable ASCII catches all of them, and that sweep is also the one that throws away every accented character in the string.

Two blank names are never the same name

Five strings that all look empty — a space, a no-break space, U+1160, U+200B and U+3164 — are five separate entries in a Set and sort in exactly that order. A uniqueness constraint on a username column will therefore not stop two accounts from both looking nameless, and a list sorted by name will order them by code point while showing five identical gaps. Percent-encoded, they are not even the same length: the fillers are nine characters (%E3%85%A4 for U+3164), a tag letter is twelve, and the no-break space is two.

Normalizing does not rescue any of it. NFKC turns U+3164 into U+1160 and U+FFA0 into U+1160; both results are still letters, still invisible, and still accepted by ^\p{L}+$. A pipeline that normalizes first and then runs a letters-only rule accepts the name at both ends, and the only thing it changed is which of the four the user ended up with.

Which one to copy

Take U+3164 if you want the one that most blank-name generators already hand out, and know that it will reserve 16 px per copy and will become U+1160 if anything normalizes your input. Take U+FFA0 if the blank name has to sit inside a tight layout, at 8 px per copy, with the same normalization caveat. Take U+1160 or U+115F if you want the value to come out of a normalizing pipeline unchanged. Take U+200B only if the thing you need is genuinely zero width — it measures 0.00 px, and it will fail any check that asks for letters.

If the reason you are reading this is the opposite problem — a name, a username or a spreadsheet column that came back with something invisible in it — the fix is not a letters-only rule. It is a decision about which characters you are willing to accept, written down as an allow list rather than a cleanup pattern, because every cleanup pattern here missed four of the seven characters in the table above.

What this page does not tell you

Whether any particular platform accepts these. What is measured here is the rules — the HTML constraint API in Chromium, the Unicode category, and the behaviour of two regex engines. Individual services layer their own checks on top, and several reject blank names outright for reasons that have nothing to do with Unicode.

The rendering columns are also specific to one engine on one machine with one font stack. The fillers painted nothing here and U+FFFC painted a box, and a font that lacks a filler entirely could draw one too. Categories, byte counts and the regex verdicts will hold anywhere; the pixel counts are a measurement of this browser, not a promise about yours.