hn.today

MicroLLM Lab – Try 7 tiny LLM's in the browser

stateofutopia.com278 points113 comments
Screenshot of MicroLLM Lab – Try 7 tiny LLM's in the browser

MicroLLM Lab is a zero-server, in-browser toolkit that runs, benchmarks, and compares tiny language models (25M-360M parameters) using hardware-accelerated WebGPU. It emphasizes on-device privacy and ultra-low latency: models are quantized to Q4 (4-bit) to shrink memory by about 75% so 100M+ models can fit in roughly 50-84 MB of browser memory, and all model files are cached privately in IndexedDB (no server calls or disk downloads required). WebGPU executes compute shaders directly on the GPU (Metal/DirectX12/Vulkan), with WASM/JS fallbacks, enabling claims like sub-10ms time-to-first-token for instant autocomplete and real-time agents while avoiding cloud costs and API billing.

The lab provides an interactive workflow - load a model, chat on-device, then run benchmark suites - measuring tokens/second (peak and sustained), suite wall time, and objective accuracy (regex/exact-token checks). Users can run sustained 256-token tests, write custom JavaScript checks evaluated against decoded text, and generate a verifiable performance certificate that records device hardware and scores for sharing. Use cases include fast triage, intent extraction, spam filtering, and local task routing. The setup is explicit that accuracy metrics are objective tests (not writing quality) and that smaller models may fail some checks, reflecting real trade-offs between size, speed, and task fit.

Read on stateofutopia.com113 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.