hn.today

Generate fonts where every LLM token is the same width

ampdot.mesh.host87 points22 comments
Screenshot of Generate fonts where every LLM token is the same width

The token-space font experiment drew mixed reactions about usability and aesthetics. Some commenters found it clever and helpful for understanding tokenization, with one saying it generated empathy for the assistant. Others reported severe performance problems: a few said Firefox hung or hit 100% CPU, while others reported smooth operation on different browser versions and operating systems. Visual fidelity also divided opinion - one person complained about horrible kerning in Safari while another reported Chrome looked fine, and someone noted punctuation centers differently than usual.

Several participants pivoted to tokenizer efficiency across languages. One provided a ranked list showing English as most token-efficient and many languages (German, Chinese, Japanese, several African and Indic languages) using substantially more tokens, arguing tokenizers are optimized for common text. Others wondered how CJK characters behave, suggesting Chinese characters may map one-to-one to tokens and thus avoid subword chopping; one commenter observed CJK text looked largely unchanged except for centered punctuation. A linked preprint was mentioned as potentially relevant, reflecting a divide between seeing tokenizers as unfair across languages and the question of whether script differences (like Han characters) truly confer efficiency.

Read on ampdot.mesh.host22 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.