Chinese model developers published a set of radical KV-cache optimizations that drastically shrink the memory needed to serve long-context models. Innovations include the MLA architecture (≈15x cache compression), Compressed Sparse Attention, Heavily Compressed Attention, and in the latest DeepSeek-V4.1-Flash, CSA2, cross-layer cache reuse, a causal encoder‑decoder layout, and FP4 caching. Together these techniques cut KV cache footprint from about 389.12 GB per million tokens (DeepSeek‑V1) to roughly 890 MB (DeepSeek‑V4.1‑Flash) - a ≈437× reduction - and intermediate versions show progressive drops (V3.2: 48.07 GB; V4‑Flash: 3.51 GB). That massive reduction targets one of the biggest costs for long-context inference: GPU VRAM occupied by the cache.
Western providers have rapidly incorporated those optimizations, quietly releasing updated models and slashing cache-related pricing. Example per‑million‑token changes: Anthropic Opus 5→5.5: input $5→$4, cache write $6.25→$5, cache read $0.50→$0.20 (≈60% cut in cache‑read cost); OpenAI GPT‑5.6→6.1 Sol: input $5→$2, cache write $6.25→$2.50, cached input $0.50→$0.10 (≈80% cut). The result is far lower inference costs, stronger inference margins for Western labs, and positive user feedback on the updated models. The dynamic flips the narrative from a simple “distillation race” to one where Chinese optimization work has become an unexpected performance lifeline for global providers.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.