Invisible username: the name renders as nothing, and the site still has to tell it apart
Short answer: an invisible username works on any platform that asks
only two questions — is the field non-empty, and is this string already taken — because
one zero-width character is not empty, is not whitespace to trim(), and is a different
value from every other blank character. It stops working the moment the platform
normalises the string before comparing it: measured across 23 invisible characters,
NFKC turns 23 distinct blank names into 18, and a trim after it erases
7 of them outright. The name you picked and the name the database stored are then
not the same name, which is why a blank username is less a trick and more a question about where
your string gets compared.
One character. Each of the first four is non-empty, is not whitespace to
trim(), and passes a field with no character-class rule. U+3164 is the one that
also passes a letters-only rule, and it reserves a full 16 px cell; U+E0020 costs two UTF-16
units, so a three-character minimum is met by two of them. U+00A0 is the control: both
trim() and \S remove it, which is why the old advice to use a
no-break space for a blank name does not survive a measurement.
The one test a username passes that a display name never does
A display name can be anything anyone else already has. A username is compared against every other one before it is accepted, and something on the platform's side decides what “already taken” means. That is the whole difference, and it is what turns a cosmetic trick into a question about what happens to your string before it is compared.
Normalisation is where the surprise lives. It is not an exotic step — it is the standard advice for any system that has to decide whether two strings are the same, and it runs before the uniqueness check on most stacks that have thought about the problem at all. The characters below are 23 invisible code points put through each normalisation form in turn.
| Form | Distinct blank names | What merges |
|---|---|---|
| NFC | 23 of 23 | nothing |
| NFD | 23 of 23 | nothing |
| NFKC | 18 of 23 | U+00A0, U+202F, U+205F, U+3000 → U+0020 U+1160, U+3164, U+FFA0 → U+1160 |
| NFKD | 18 of 23 | the same two groups |
Five of the blank names you could have chosen are not five names to a uniqueness check that runs after NFKC. Four of them become an ordinary space, and three of them — the two Hangul fillers and the halfwidth filler — become the same Hangul filler. If you pick U+3164 and somebody else already has U+1160, the check is not wrong about the collision. It is working from a shorter list than the one you picked from.
Trim and normalise: the order makes no difference, which is the part that feels wrong
Trimming is the obvious repair, and it is the step that actually removes them. Measured on all 23
characters in JavaScript, trim() empties exactly seven: U+00A0, U+2028, U+2029,
U+202F, U+205F, U+3000 and U+FEFF. Run the normalisation first and then the trim, and it empties
the same seven. Run the trim first and then the normalisation, and it is still the same seven. The
order changes nothing about which blank usernames survive, which is the opposite of what you would
expect from two steps that each rewrite the string.
Python agrees on the count with one exception, and the exception is worth knowing. Python's
strip() empties six of the 23, because it does not treat U+FEFF as whitespace:
isspace() is False for it and \s does not match it. So a
username consisting of a single zero-width no-break space is empty to a browser running
trim() and one character long to a Python service. Validation split across a
browser and a server will disagree about that name, and neither side is broken.
| Step | Emptied | Which characters |
|---|---|---|
JavaScript trim() | 7 of 23 | U+00A0, U+2028, U+2029, U+202F, U+205F, U+3000, U+FEFF |
Python strip() | 6 of 23 | the same list without U+FEFF |
| NFKC then trim | 7 of 23 | identical to the trim alone |
| trim then NFKC | 7 of 23 | identical again |
Appending one to a real name beats exact matching, not substring matching
The other half of the problem is a name that is not blank, but is not the name it looks like
either. admin followed by a zero-width space fails === against
admin, passes startsWith and includes, and is reported as
equal by localeCompare. A blocklist built on substring matching
catches the imitation; a check built on exact matching lets it through. Two of the most ordinary
ways of asking “are these the same name” give opposite answers about the same pair of
strings.
Appended to admin | === | includes | localeCompare |
|---|---|---|---|
| U+200B | false | true | 0 — equal |
| U+200C | false | true | 0 — equal |
| U+200D | false | true | 0 — equal |
| U+2060 | false | true | 0 — equal |
| U+FEFF | false | true | 0 — equal |
| U+00AD | false | true | 0 — equal |
| U+180E | false | true | 0 — equal |
| U+E0020 | false | true | 0 — equal |
| U+3164 | false | true | not equal |
| U+FFA0 | false | true | not equal |
| U+00A0 | false | true | not equal |
| U+3000 | false | true | not equal |
The four at the bottom are the interesting ones, because the collator is the only thing in the table that notices them. They are also the four that a trim will take off, so the pair of behaviours line up: the characters a collator treats as a different name are the characters a trim deletes, and the characters a collator treats as the same name are the ones that survive it.
Lowercasing does not help either. Many platforms fold case before the uniqueness check, and
measured in Node, toLowerCase() and toUpperCase() both return all of
these unchanged. There is no case to fold, so the step that was supposed to catch
“Admin” against “admin” passes a blank name straight through.
In the profile URL a blank name is not blank
The profile URL is where the trick stops being invisible. Parsing
https://example.com/u/<name> with a trailing ordinary space in Node drops the
space — the path comes back as /u/ and nothing else. Every invisible character
is kept, and written out as percent escapes instead. A name that renders as no ink at all becomes
six to twelve visible characters in the address bar.
| Character | In the path | Visible characters |
|---|---|---|
| ordinary space | stripped by the parser | 0 |
| U+00A0 | %C2%A0 | 6 |
| U+200B | %E2%80%8B | 9 |
| U+3164 | %E3%85%A4 | 9 |
| U+2060 | %E2%81%A0 | 9 |
| U+E0020 | %F3%A0%80%A0 | 12 |
Two consequences follow. A link to a blank profile cannot be shortened by eye and cannot be retyped from a screenshot, because what is in the address bar is not what is on the page. And the name is not blank in a server log either: the request line carries the escapes, so a blank name shows up in an access log as a row of percent signs where every other row has a word.
What the form itself accepts
Before any of that, the field has to let the character in. Measured in Chromium by putting each
of the 23 characters into a real <input> and reading
validity.patternMismatch, the pattern attribute is the whole gate — and the
pattern most username fields actually use refuses everything.
| Rule on the field | Accepts | What it lets through |
|---|---|---|
pattern="[A-Za-z0-9_]+" | 0 of 23 | nothing |
pattern="\w+" | 0 of 23 | nothing — \w is ASCII here |
pattern="\S+" | 16 of 23 | everything except the 7 JavaScript counts as whitespace |
pattern="\p{L}+" | 4 of 23 | U+115F, U+1160, U+3164, U+FFA0 |
| no pattern at all | 21 of 23 | everything except U+2028 and U+2029 |
The letters-only row is the one that catches people out. Loosening a rule to
\p{L} is the standard advice for accepting names in any alphabet, and it admits all
four Hangul fillers while continuing to refuse the zero-width space that the same advice-giver
would probably have expected it to block. Nothing in the middle works either: the four fillers are
letters, so any rule that accepts letters accepts them.
The length rules are worth knowing separately, because they count the wrong thing. A
required field treats a single invisible character as filled in — measured, all
23 of 23, including the ones whose trim() length is zero — while an empty
string fails. And minlength counts UTF-16 units rather than characters: typed into a
field with a minimum of three, two zero-width spaces measure 2 and are rejected as too short,
while two tag characters measure 4 and pass, even though both are two keystrokes and neither shows
anything. At a maximum of five the same asymmetry appears from the other side: five zero-width
spaces fit, and only two tag characters do.
Where the advice splits
Everybody who has worked on this agrees on the first move: normalise the string before you compare it, and do it in one place. The standard illustration is a music service in 2013, where two usernames that were different at the Unicode level normalised into the same form and one account could take over the other. The measured table above is that incident in miniature — five distinct inputs, two normalised outputs.
The second position is that normalisation is being applied at the wrong moment rather than in the
wrong way. On that reading the check that stops somebody registering both
admin and Admin only has to run once, when the account is created; at
login the server should not accept Admin as a valid way to reach the account
admin at all. If you take that seriously, a blank username stops being a Unicode
problem and becomes a decision about when the comparison happens — and the measurements here
say the two moments need different rules, because the same string is a collision at signup and a
login failure afterwards.
The third position is the minority one, and it comes in two directions that both cut against the
other two. One argument is that the real risk is the dependency rather than the character: a
platform that accepts Unicode usernames has to pull a large Unicode library into the
authentication path, which is a lot of surface area for a name. The other is that restricting
usernames to ASCII is not the safe option either, because a name is a string that gets displayed
and stored, and you do not want a bell or a null in it any more than you want a zero-width space.
Neither side of that argument is settled by anything measured on this page. What the measurements
do say is that the characters which survive every gate are the ones nobody thought to list: the
four Hangul fillers pass a letters-only rule, and the two rules that catch them —
trim() and a collator — are the two most people would not think to run on a
username at all.
What this page does not tell you
Whether any particular platform accepts these. What is measured here is the rules — the Unicode normalisation forms, the HTML constraint API in Chromium, and the behaviour of two language runtimes. Individual services layer their own checks on top, and several reject blank names outright for reasons that have nothing to do with Unicode. A platform that runs none of the steps in the tables above will take a blank username; a platform that runs all of them will not, and neither outcome says anything about the other.
The set of 23 is also the set measured, not a canonical list of every invisible character.
A different set moves the counts, though not the direction of any result: more characters mean
more merges under NFKC and more survivors of a trim. The runtimes are single builds as well —
Python 3.13.14 carries Unicode 15.1.0 and Node 22.22.2 carries Unicode 17.0, and a platform on a
third version could normalise a code point differently from both. Categories, byte counts and the
=== verdicts will hold anywhere; the exact merge list is a measurement of these two
versions.
Measured 2026-10-04 with Chromium via Playwright, Python 3.13.14 (unicodedata 15.1.0) and Node 22.22.2 (ICU 78.2, Unicode 17.0). Related: the full list of invisible characters, the four that are letters and print nothing, why there is no empty character, how to find hidden characters in a file, comparing two texts that look identical, what a browser does with characters you cannot see, tag characters that hide a whole ASCII message, and the zero-width space remover.