hn.today

ESP32S3 cluster running 1.58-bit (BitNet) Language model

github.com150 points31 comments
Screenshot of ESP32S3 cluster running 1.58-bit (BitNet) Language model

This project implements a distributed inference engine that runs a sliced 0.5B language model across seven ESP32-S3 microcontrollers using 1.58-bit (BitNet) ternary quantization. One board is the master that handles prompt I/O, a BPE tokenizer, INT4 token embeddings (~14 MB in flash), the final RMS norm and the LM head with greedy sampling. Six compute nodes form a SPI daisy-chain and each runs a block of transformer layers (24 layers total divided into 4-layer chunks per node), executing RMSNorm (FP16 scaled to FP32), 1.58-bit attention (Q/K/V/O projections with RoPE) and 1.58-bit MLPs. KV caches live in PSRAM and data flows as FP32 hidden-state vectors over dual SPI channels for low-latency layer chaining.

The repository includes ESP-IDF firmware for master and compute nodes (C++ and assembly-optimized MAC kernels), a Python toolchain for quantization-aware training, model slicing, tokenizer packing and embedding packing, and wiring/flash workflow instructions. Key implementation details: bitlinear ternary linear layers with hand-optimized assembly, LUT acceleration, partitioned binaries for flash alignment, and dual-channel SPI DMA for daisy-chain communication. The project is MIT-licensed, documented with diagrams and hardware photos, and aims to demonstrate practical microcontroller-scale distributed LLM inference using extreme quantization.

Read on github.com31 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.