hn.today

AI is now capable of developing its own inference hardware

github.com168 points220 comments
Screenshot of AI is now capable of developing its own inference hardware

openTPU is an open-source AI accelerator project that packages a full stack - SystemVerilog RTL, an 8x32-bit-word ISA and bit-exact Python simulator, a kernel language and compiler, and host PCIe software - into a single monorepo so readers can follow a model from a Python matmul down to wires. Framed as an experiment in how far AI agents can push hardware design, the implementation runs on an Inspur YPCB-00338 FPGA card (Xilinx Kintex-7 xc7k480t, two DDR3 channels) and reproduces simulator tokens bit-for-bit. Measured decode speeds span dozens of models: small LMs like LFM2.5-230M reach ~59-86 tok/s (int8 to 4-bit), mid models like Qwen3.5-2B hit ~12 tok/s, and MoE setups run larger networks by streaming expert weights from host storage (e.g., LFM2.5-8B at 10.6 tok/s with 98.5% slot hits; QWEN3.5-35B at 3.95 tok/s with 153 MB/token streamed). DRAM utilization is high (typically 82-94% of DDR3 peak).

The architecture deliberately keeps everything explicit: a simple sequencer issues one instruction per cycle to DMA, a systolic matrix unit, a vector FP32 unit and a quantizer; there is no cache and all data movement is visible in traces. Quantization uses 4-bit FP4 with two-level block scales (≈4.25 bits/weight) while keeping the LM head in int8, giving ~40-45% decode speedups at a measurable perplexity cost. Host involvement is minimal - most decodes run a compiled device program with logits streamed back - and the repo documents calibration, performance tools, kernel examples and how offload/streaming for large MoE models is handled.

Read on github.com220 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

Nano Banana 2.1

Nano Banana 2.1

Google AI Studio announced Nano Banana 2.1, an improved image generation model with better visual design, mask-based editing, and natural-looking images. The model outperforms previous versions across all metrics and is available for testing at ai.studio. (twitter.com)

EmbeddingGemma 2

EmbeddingGemma 2

EmbeddingGemma 2 is a new multimodal embedding model developed by Google, capable of handling both text and vision data. It offers a moderate size of 270 million parameters for text and 440 million for combined text and vision, aiming to improve how large language models and AI agents work. (blog.google)

What Is Codemode

What Is Codemode

Codemode introduces a way to integrate tools directly into the execution environment for language models, allowing more complex and native interactions. It emphasizes the separation between the trusted harness and the target environment where tools run, enabling better security and functionality. (lucumr.pocoo.org)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.