Invisible letter: four of them are real letters, and every letters-only rule waves them through
Short answer: of the twelve code points Unicode names as fillers,
exactly four are letters — U+3164 HANGUL FILLER, U+1160 HANGUL JUNGSEONG FILLER,
U+115F HANGUL CHOSEONG FILLER and U+FFA0 HALFWIDTH HANGUL FILLER. They print nothing and
they are not whitespace, which is the combination that makes them useful: \p{L}
accepts them, isalpha() accepts them, and no trim will delete them. The zero-width
space you were handed instead is none of those things — it is a format character, it fails
every letters-only test, and it measures 0.00 px where a filler reserves a whole character cell.
One character, three bytes as UTF-8, no ink on screen. U+3164, U+1160 and U+115F each advance
the cursor by 16 px in a 16 px font; U+FFA0 advances it by 8 px, so it is the one to use where
a blank name is pushing a layout around. All four pass ^\p{L}+$. None of them is
removed by trim().
The four, measured
Twelve Unicode characters carry the word FILLER in their name. Four are letters, seven are punctuation and one is a combining mark. Only the four letters behave like a character you typed: they occupy a cell, they survive a trim, and they answer true to every “is this a letter” question anyone thinks to ask.
| Code point | Name | Cat. | Width / char | Ink painted | Bytes | NFKC | JS variable name |
|---|---|---|---|---|---|---|---|
| U+3164 | HANGUL FILLER | Lo | 16.00 px | 0 px | 3 | U+1160 | legal |
| U+1160 | HANGUL JUNGSEONG FILLER | Lo | 16.00 px | 0 px | 3 | unchanged | legal |
| U+115F | HANGUL CHOSEONG FILLER | Lo | 16.00 px | 0 px | 3 | unchanged | legal |
| U+FFA0 | HALFWIDTH HANGUL FILLER | Lo | 8.00 px | 0 px | 3 | U+1160 | legal |
| U+200B | ZERO WIDTH SPACE | Cf | 0.00 px | 0 px | 3 | unchanged | syntax error |
| U+E0054 | TAG LATIN CAPITAL LETTER T | Cf | 0.00 px | 0 px | 4 | unchanged | syntax error |
| U+FFFC | OBJECT REPLACEMENT CHARACTER | So | 16.00 px | 1070 px | 3 | unchanged | syntax error |
The ink column is the one nobody publishes. It is a pixel count of the rendered glyph against a blank region of the same size: ten copies of U+3164 in a 16 px font measured 160 px wide and painted zero non-white pixels. The character reserves a full cell and puts nothing in it. That is the whole trick, and it is also the tell — anything with a background, a border or an underline behind a blank name will show a 16 px gap where the eye expects none.
U+FFFC is in the table because it is the usual next guess and it fails. It is the character meant to stand in for a missing object, it is printable, and in Chromium it paints a visible box: 1070 non-white pixels for ten copies, against 0 for the fillers and 492 for a Cyrillic а. Counting painted pixels is the only test that separates the four real invisible letters from an ordinary glyph your font happens to lack.
The letters-only rule is the hole, not the fix
The advice given whenever somebody asks how to accept names in any alphabet is to match
\p{L}. It is in the accepted answers, it is the obvious reading of “any
letter”, and it lets all four of these in. Measured in Chromium with a real form, a field
carrying pattern="\p{L}+" validates a value of U+3164, U+1160, U+115F or U+FFA0, and
rejects U+200B, U+00A0, U+FEFF and the tag characters. A field carrying
pattern="[A-Za-z]+" rejects everything invisible — and every name that is not
written in the Latin alphabet with it.
That is the fork, and there is no clever third option in the middle of it. Either your rule
accepts exactly the characters somebody can type on a US keyboard, or it admits every character
Unicode files under Letter, including four that nobody can see. The rules that look like a
compromise are not: [\p{L}\p{M}] accepts the four as well, because the marks were
never the problem.
One rule, two languages, opposite answers
Write the same check in Python and in JavaScript and you get different verdicts on the same string, because the two languages do not mean the same thing by the same pattern.
| Rule | U+3164 | U+1160 | U+115F | U+FFA0 | U+200B | U+E0054 |
|---|---|---|---|---|---|---|
Python re.fullmatch(r"\w+") | pass | pass | pass | pass | fail | fail |
JavaScript /^\w+$/ | fail | fail | fail | fail | fail | fail |
Python s.isalpha() | pass | pass | pass | pass | fail | fail |
JavaScript /^\p{L}+$/u | pass | pass | pass | pass | fail | fail |
Python s.isprintable() | pass | pass | pass | pass | fail | fail |
Python s.isidentifier() | pass | pass | pass | pass | fail | fail |
either language [A-Za-z0-9_]+ | fail | fail | fail | fail | fail | fail |
Python’s \w is Unicode-aware and matches any letter, so a blank name is a
word. JavaScript’s \w is [A-Za-z0-9_] and nothing else, so the
same blank name is not. Both are documented, both are correct, and a stack that validates in the
browser with \w and again on the server with \w will reject the name in
one place and accept it in the other.
The identifier row is the sharpest version of the same thing. All four fillers are legal
JavaScript variable names on their own — var followed by U+3164 parses. It has
been called a bug in the identifier rules and it has been defended as the price of letting people
write code in their own script. Either way it is the property being relied on here: the character
is a letter to the language, not to the reader.
The characters named like letters are not letters
The tag block, U+E0020 to U+E007F, is where the naming goes the other way. Its characters are
called TAG LATIN CAPITAL LETTER T, TAG LATIN SMALL LETTER A and so on down the ASCII range, and
they are not letters. Their category is Cf. \p{L} does not match them. They cannot
start or continue a variable name. They cost four bytes and two UTF-16 code units each, against
three bytes and one unit for a filler, and they do something no filler does: they attach to the
character in front of them.
Measured with the grapheme segmenter in Chromium, the string a + U+E0054 +
b breaks into two clusters, not three — the tag character is part of the
a. The same test on U+3164 gives three. So a single backspace after a visible letter
with a tag character behind it takes the letter with it, and a counter that reports
“characters” by counting clusters reports one fewer than a counter that counts code
points. The fillers are ordinary standalone characters; the tag characters are passengers.
That difference also decides what your cleanup finds. The popular one-liner that strips
[\u200B-\u200D\uFEFF] catches the zero-width space and the byte-order mark and
misses all four invisible letters, because letters are not in its range. \p{C} and
\p{Cf} catch the tag characters and the zero-width family and miss all four
invisible letters, for the same reason in reverse. Only a sweep for anything outside printable
ASCII catches all of them, and that sweep is also the one that throws away every accented
character in the string.
Two blank names are never the same name
Five strings that all look empty — a space, a no-break space, U+1160, U+200B and
U+3164 — are five separate entries in a Set and sort in exactly that order. A uniqueness
constraint on a username column will therefore not stop two accounts from both looking nameless,
and a list sorted by name will order them by code point while showing five identical gaps.
Percent-encoded, they are not even the same length: the fillers are nine characters
(%E3%85%A4 for U+3164), a tag letter is twelve, and the no-break space is two.
Normalizing does not rescue any of it. NFKC turns U+3164 into U+1160 and U+FFA0 into U+1160;
both results are still letters, still invisible, and still accepted by ^\p{L}+$. A
pipeline that normalizes first and then runs a letters-only rule accepts the name at both ends,
and the only thing it changed is which of the four the user ended up with.
Which one to copy
Take U+3164 if you want the one that most blank-name generators already hand out, and know that it will reserve 16 px per copy and will become U+1160 if anything normalizes your input. Take U+FFA0 if the blank name has to sit inside a tight layout, at 8 px per copy, with the same normalization caveat. Take U+1160 or U+115F if you want the value to come out of a normalizing pipeline unchanged. Take U+200B only if the thing you need is genuinely zero width — it measures 0.00 px, and it will fail any check that asks for letters.
If the reason you are reading this is the opposite problem — a name, a username or a spreadsheet column that came back with something invisible in it — the fix is not a letters-only rule. It is a decision about which characters you are willing to accept, written down as an allow list rather than a cleanup pattern, because every cleanup pattern here missed four of the seven characters in the table above.
What this page does not tell you
Whether any particular platform accepts these. What is measured here is the rules — the HTML constraint API in Chromium, the Unicode category, and the behaviour of two regex engines. Individual services layer their own checks on top, and several reject blank names outright for reasons that have nothing to do with Unicode.
The rendering columns are also specific to one engine on one machine with one font stack. The fillers painted nothing here and U+FFFC painted a box, and a font that lacks a filler entirely could draw one too. Categories, byte counts and the regex verdicts will hold anywhere; the pixel counts are a measurement of this browser, not a promise about yours.
Measured 2026-09-29 with Chromium via Playwright, Python 3.13.14 (unicodedata 15.1.0) and Node 22.22.2 (ICU, Unicode 17.0). Related: the full list of invisible characters, why there is no empty character, how to find hidden characters in a file, what one does to a form field, tag characters that hide a whole ASCII message, what a browser does with characters you cannot see, the zero-width space remover, and who can still read text after you hide it.