Tech & AI

Tech & AI

How software, search engines and language models actually handle text - and what that means for anyone who writes or ships it.

Why this section exists. Styled Unicode looks like a design choice but is an encoding choice, and every system that reads text downstream notices. This section follows that thread into the technology.

Every generator on this site produces characters, not formatting. That single fact is why styled text survives a paste into a bio - and why it costs a language model ten times the tokens, why a screen reader may spell it out letter by letter, and why a spoofed domain can look identical to the real one. This section follows those consequences into the systems themselves.

The articles here are written for people who build or operate software that handles text, and for writers who want to know what happens to their words after they leave the page. Claims are measured or sourced; where we ran a test, the method is in the article so you can repeat it.

Topics

Encoding & Unicode

Codepoints, normalisation, byte cost, and the standards that decide what a character is.

0 articles

Language models & search

Tokenization, indexing, answer engines, and why machines read characters rather than shapes.

1 article

Accessibility technology

Screen readers, assistive tooling, and what styled text does to them.

0 articles

Text & security

Homoglyphs, confusables, spoofed identifiers, and the defences built around them.

0 articles

Latest

What Styled Unicode Costs a Language Model

The same eleven characters cost 2 tokens in plain text and 21 to 30 tokens in every decorative Unicode style we measured. An 80-character bio went from 20 tokens to 188. The mechanism is simple, and so is the fix.

Read the guide →
Write for this section. Tech & AI accepts guest and sponsored contributions under a published set of standards and a link policy. Pitches that fit the topics above are read; pitches that do not are declined.