Contemporary large language models rely on subword tokenization, which obscures character-level information important for tasks like code and biological sequences. A general method called byteification retrofits existing subword models into byte-level models with minimal extra pretraining (under 1% of a typical budget, 49.1B tokens), producing byteified models such as Bolmo 7B/1B, Bwen 8B and Blama 8B from open subword parents. The architecture is a latent tokenizer language model (LTLM) that first encodes bytes locally, predicts non-causal patch boundaries to aggregate bytes into patches, pools patch representations, processes them with a global transformer, and depoools back to bytes; shallow wide local models using mLSTM layers handle byte-level contextualization. The two-stage conversion and architectural choices resolve expressivity mismatches between subword and latent patch tokenizations while avoiding training from scratch.
Empirically, byteified models substantially outperform prior publicly available byte-level LLMs of comparable size and recover or exceed aspects of their subword sources: Bolmo 7B shows a +16.5% absolute gain on STEM tasks over a prior byte model and improves character understanding and some coding tasks relative to its source Olmo 3. Byteified models allow faster inference by increasing bytes-per-patch and can reuse existing ecosystem components for adaptation without extra training. These results remove a major performance barrier to end-to-end byte-level modeling, promising better fine-grained textual understanding, reduced English-centric tokenization bias and gains in computational efficiency for scientific and technical domains.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.