hn.today

Show HN: Long term Memory and 50M token window for LLM

huggingface.co5 points0 comments
Screenshot of Show HN: Long term Memory and 50M token window for LLM

This work presents galahad-kv, a memory layer that persists a transformer model's internal key-value (KV) states for blocks of roughly 16,000 tokens to encrypted local NVMe and later reloads them byte-exact without recomputation. The system was integrated with vLLM and evaluated on a 50 million token stream served from a single NVIDIA H100 using Gemma 4 12B and 31B models. The test protocol is designed to resist common benchmark gaming and supports a single-GPU reproduction with public software and a permissive license; project code and a wiki are provided.

Results show perfect reuse of stored blocks in probing (100/100 blocks loaded, depths 0-50M) on both models. Loading persisted KV state was 2.8-4.3× faster than recomputing and consumed 8.8-12.3× less GPU energy, while GPU memory usage stayed flat across the 50M-token stream. In factual retrieval tests where facts were planted millions of tokens earlier, the 12B model answered correctly 82/100 times and the 31B 98/100 times; neither model hallucinated answers. Important constraints: this approach reuses one loaded block at a time (it is not an expanded attention window), writing memory is a one-time cost, and the on-disk store requires terabytes of NVMe.

Read on huggingface.co0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.