Tech & AI
How software, search engines and language models actually handle text - and what that means for anyone who writes or ships it.
Every generator on this site produces characters, not formatting. That single fact is why styled text survives a paste into a bio - and why it costs a language model ten times the tokens, why a screen reader may spell it out letter by letter, and why a spoofed domain can look identical to the real one. This section follows those consequences into the systems themselves.
The articles here are written for people who build or operate software that handles text, and for writers who want to know what happens to their words after they leave the page. Claims are measured or sourced; where we ran a test, the method is in the article so you can repeat it.
Topics
Encoding & Unicode
Codepoints, normalisation, byte cost, and the standards that decide what a character is.
0 articlesLanguage models & search
Tokenization, indexing, answer engines, and why machines read characters rather than shapes.
1 articleAccessibility technology
Screen readers, assistive tooling, and what styled text does to them.
0 articlesText & security
Homoglyphs, confusables, spoofed identifiers, and the defences built around them.
0 articles