This project implements a distributed inference engine that runs a sliced 0.5B language model across seven ESP32-S3 microcontrollers using 1.58-bit (BitNet) ternary quantization. One board is the master that handles prompt I/O, a BPE tokenizer, INT4 token embeddings (~14 MB in flash), the final RMS norm and the LM head with greedy sampling. Six compute nodes form a SPI daisy-chain and each runs a block of transformer layers (24 layers total divided into 4-layer chunks per node), executing RMSNorm (FP16 scaled to FP32), 1.58-bit attention (Q/K/V/O projections with RoPE) and 1.58-bit MLPs. KV caches live in PSRAM and data flows as FP32 hidden-state vectors over dual SPI channels for low-latency layer chaining.
The repository includes ESP-IDF firmware for master and compute nodes (C++ and assembly-optimized MAC kernels), a Python toolchain for quantization-aware training, model slicing, tokenizer packing and embedding packing, and wiring/flash workflow instructions. Key implementation details: bitlinear ternary linear layers with hand-optimized assembly, LUT acceleration, partitioned binaries for flash alignment, and dual-channel SPI DMA for daisy-chain communication. The project is MIT-licensed, documented with diagrams and hardware photos, and aims to demonstrate practical microcontroller-scale distributed LLM inference using extreme quantization.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.