openTPU is an open-source AI accelerator project that packages a full stack - SystemVerilog RTL, an 8x32-bit-word ISA and bit-exact Python simulator, a kernel language and compiler, and host PCIe software - into a single monorepo so readers can follow a model from a Python matmul down to wires. Framed as an experiment in how far AI agents can push hardware design, the implementation runs on an Inspur YPCB-00338 FPGA card (Xilinx Kintex-7 xc7k480t, two DDR3 channels) and reproduces simulator tokens bit-for-bit. Measured decode speeds span dozens of models: small LMs like LFM2.5-230M reach ~59-86 tok/s (int8 to 4-bit), mid models like Qwen3.5-2B hit ~12 tok/s, and MoE setups run larger networks by streaming expert weights from host storage (e.g., LFM2.5-8B at 10.6 tok/s with 98.5% slot hits; QWEN3.5-35B at 3.95 tok/s with 153 MB/token streamed). DRAM utilization is high (typically 82-94% of DDR3 peak).
The architecture deliberately keeps everything explicit: a simple sequencer issues one instruction per cycle to DMA, a systolic matrix unit, a vector FP32 unit and a quantizer; there is no cache and all data movement is visible in traces. Quantization uses 4-bit FP4 with two-level block scales (≈4.25 bits/weight) while keeping the LM head in int8, giving ~40-45% decode speedups at a measurable perplexity cost. Host involvement is minimal - most decodes run a compiled device program with logits streamed back - and the repo documents calibration, performance tools, kernel examples and how offload/streaming for large MoE models is handled.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.