What Styled Unicode Costs a Language Model
The same eleven characters cost 2 tokens in plain text and 21 to 30 tokens in every decorative Unicode style we measured. An 80-character bio went from 20 tokens to 188. The mechanism is simple, and so is the fix.
In short
Decorative Unicode text - mathematical bold, script, fraktur, circled letters, fullwidth - costs a language model roughly ten to fifteen times more tokens than the same words in plain ASCII, and glitch (zalgo) text over sixty times more. Measured with the GPT-4 (cl100k_base) and GPT-4o (o200k_base) tokenizers on 29 September 2026: "Hello world" is 2 tokens plain and 21-30 tokens styled. The cause is that tokenizers learn merges from frequent byte sequences and these characters are rare three- and four-byte sequences, so they fall back to byte-level pieces. Unicode NFKC normalisation folds most styled alphabets back to ASCII before tokenizing.
- "Hello world": 2 tokens plain; 19-30 tokens in every styled alphabet tested; 103-133 tokens as glitch text.
- A realistic 80-character bio: 20 tokens plain, 122 in fullwidth, 188 in bold script.
- Cause: styled characters are 3-4 bytes each and rare in training data, so byte-pair tokenizers cannot merge them into whole-word tokens.
- NFKC normalisation maps mathematical, enclosed and fullwidth letters back to ASCII; it does not fix small caps, pseudo-alphabets or combining marks.
- It matters wherever text is priced, chunked or embedded: prompts, retrieval pipelines, logs, moderation queues.
The measurement
We took a short string, rendered it in the styles this site generates, and counted tokens with two production tokenizers: cl100k_base, used by GPT-4 and GPT-3.5, and o200k_base, used by GPT-4o. The library was js-tiktoken 1.x, the date 29 September 2026. Byte counts are UTF-8.
| Style | Characters | UTF-8 bytes | Tokens (GPT-4) | Tokens (GPT-4o) |
|---|---|---|---|---|
| Plain ASCII | 11 | 11 | 2 | 2 |
| Bold serif | 11 | 41 | 30 | 21 |
| Script | 11 | 37 | 26 | 27 |
| Fraktur | 11 | 40 | 29 | 30 |
| Circled | 11 | 31 | 30 | 21 |
| Small caps | 11 | 26 | 25 | 21 |
| Fullwidth | 11 | 31 | 21 | 19 |
| Glitch, medium | 60 | 109 | 107 | 103 |
| Glitch, heavy | 73 | 135 | 133 | 127 |
The string was "Hello world". Every styled alphabet landed between ten and fifteen times the plain count. Glitch text - which stacks combining marks on each letter - was fifty to sixty-five times. The two tokenizers disagree on details but not on the shape of the result.
A more realistic case: an eighty-character Instagram-style bio ("Photographer based in Lisbon. Weddings, portraits, film. Bookings open for 2027.") is 20 tokens in plain text with o200k_base, 122 in fullwidth, and 188 in bold script. If that bio is one of thousands being summarised, classified or embedded, the styled versions cost six to nine times more for every pass.
Why this happens
Byte-pair encoding starts from single bytes and repeatedly merges the pair of adjacent tokens that appears most often in its training corpus. The vocabulary that results is a record of what was common. "Hello" and " world" were common, so each became one token.
Mathematical bold, script and fraktur letters live in Unicode's Supplementary Multilingual Plane, so each is four bytes in UTF-8. Circled letters and fullwidth forms are three bytes. None of them were common in the corpus. With no learned merges to apply, the tokenizer falls back to byte-level pieces - and four bytes of an unfamiliar character can come out as two or three tokens.
Glitch text is the extreme case because there is more of it: each visible letter carries several combining diacritics, every one of which is its own two-byte character with the same problem.
Where the cost shows up
- **Prompts and context windows.** User-generated text pasted into a prompt - bios, comments, product titles - can consume ten times the budget it appears to.
- **Retrieval and chunking.** Chunk sizes are measured in tokens. A styled passage produces fewer characters per chunk and a worse embedding, because the model has seen the plain spelling far more often than the decorated one.
- **Moderation and classification.** Decorated spellings are a known way to slip past keyword filters. A model-based classifier is more robust, but pays the token cost on every item.
- **Logs and analytics.** Anything that stores or bills by token inherits the multiplier silently.
The fix: normalise before you tokenize
Unicode defines compatibility decompositions for characters that are stylistic variants of others. NFKC normalisation applies them: mathematical bold 𝐇 becomes H, circled ⓗ becomes h, fullwidth H becomes H. One call in most languages - normalize("NFKC") in JavaScript, unicodedata.normalize in Python - collapses the mathematical alphabets, enclosed letters and fullwidth forms back to ASCII before the tokenizer sees them.
It is not complete. Small-capital letters are phonetic characters, not compatibility variants, so ᴀ stays ᴀ. Pseudo-alphabets that borrow lookalikes from Cyrillic, Cherokee or Canadian syllabics are different letters entirely and stay as they are. Combining marks are not removed by NFKC; stripping glitch text means decomposing with NFKD and removing characters in the Mark category. For the confusable cases there is the Unicode confusables data, which is the same table used to detect spoofed domain names.
- Apply NFKC to all user-supplied text at the boundary, before it reaches a prompt, an index or a log.
- Decompose with NFKD and drop combining marks (Unicode category M) if you need to neutralise glitch text.
- Map confusables with the Unicode confusables table if you are matching identifiers or filtering.
- Keep the original alongside the normalised form if you ever need to display it as the user wrote it.
For anyone writing rather than building: this is the machine-side reason the guides here keep saying to leave important words in plain letters. A search engine, an assistant summarising your page and a moderation model all read codepoints, not shapes.
Reproducing this
Install js-tiktoken, call getEncoding("o200k_base").encode(text).length on the plain and styled strings, and compare. The styled strings can be produced with any of the generators on this site; the bold text generator and glitch text generator cover the two ends of the range. Counts will drift slightly between library versions; the ratios will not.
Frequently asked questions
How many tokens does fancy Unicode text use?
In our measurement, ten to fifteen times the plain-text count for styled alphabets (bold, script, fraktur, circled, fullwidth) and over sixty times for glitch text, using the GPT-4 and GPT-4o tokenizers.
Why do tokenizers handle styled letters so badly?
Byte-pair encoding learns merges from byte sequences that were common in training data. Mathematical and enclosed letters are 3-4 bytes each and rare, so no merges exist for them and each character is split into several byte-level tokens.
Does NFKC normalisation fix it?
For the mathematical alphabets, enclosed letters and fullwidth forms, yes: NFKC maps them to ASCII. It does not map small caps or pseudo-alphabets borrowed from other scripts, and it leaves combining marks in place; those need separate handling.
Does this affect how well a model understands the text?
Often, yes. A model has seen the plain word far more than any styled spelling, so styled text can degrade instruction-following and retrieval matching as well as costing more. Normalising input before the model sees it avoids both problems.
Is this specific to OpenAI tokenizers?
The exact counts are. The mechanism - rare multi-byte characters falling back to byte pieces - applies to any byte-level BPE tokenizer, which covers most current models.