A Self-Pruning Transformer introduces Universal Attention, a trainable attention architecture that dramatically compresses KV-caches by learning composite decay mechanisms while preserving expressive RoPE positional embeddings and Softmax attention. It addresses the deployment bottleneck from large KV-caches by replacing simple decay heuristics used in prior work - heuristics that often collapse into sliding-window eviction - with a richer, complementary set of decay functions that model complex key statistics and interactions. The learned composite decay naturally behaves as an adaptive pruning criterion, identifying and removing tokens that contribute least to attention computation rather than relying on fixed-distance eviction.
Experiments show extreme KV-cache compression with practical benefits: Universal Attention achieves state-of-the-art 10× compression on natural language and synthetic tasks while improving downstream performance compared to leading compression baselines and even unpruned oracles. It also exhibits superior long-context generalization, reporting up to 25× compression at length 16k without degrading accuracy. The method is end-to-end trainable and designed to deliver lower memory footprints and faster inference without sacrificing the positional expressivity and Softmax dynamics central to transformer performance.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.