hn.today

Font where each token is equal-width

twitter.com26 points14 comments
Screenshot of Font where each token is equal-width

A designer is building a prototype font in which every token defined by a chosen tokenizer occupies equal horizontal space, making the invisible tokenization that models use visible in regular text. A web tool lets users combine any base font with a tokenizer to generate “token-space” fonts and install them in chat apps so AI outputs can be read exactly as the model segments and lays out tokens. The stated goal is practical: exposing token boundaries should make it easier to diagnose why models exhibit particular quirks and failure modes by letting people see tokenization patterns at a glance.

Testing shows most tokenizer-font pairings map accurately (>99% for many), but some modern models (notably GLM-5 and LLaMa 3) still produce unreliable mappings and there are recurring edge cases around merge rules, text wrapping, Unicode handling, and kerning. Implementation required workarounds because OpenType limits ligatures and many tokenizers behave oddly; the solution uses GSUB and cmap tables, and for some tokenizers (e.g., Claude) the font performs dynamic-programming/shortest-path searches to create compact glyph sequences. Compact fonts also compress better, and some combinations (SF Compact + large GPT vocabulary) remain quite readable.

Read on twitter.com14 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.