hn.today

Why isn't the industry freaking out about DeepSeek 4.1 Flash?

dgt.is380 points320 comments
Screenshot of Why isn't the industry freaking out about DeepSeek 4.1 Flash?

A heavy user reports that DeepSeek 4.1 Flash, after a month of intensive use across a dozen projects, delivers frontier-like performance at a tiny fraction of the cost. In real sessions the model is indistinguishable from Opus for conversation, coding, speed, and complex planning; with an OpenCode Go $10/month subscription the user runs long, exploratory workflows for pennies, often under $1 per all-day session and as little as $0.003 per simple task. The workflow shift is profound: cheap, high-quality unattended tasks, exploratory UI testing, and iterative development become routine, with occasional calls to Opus 5.5 or GLM only for final code reviews or fresh perspectives.

The technical lever behind this is a roughly 437x reduction in KV cache size versus DeepSeek V1, dramatically cutting GPU memory footprint, cost, and environmental impact, and contributing efficiency gains across other models. Those cache optimizations make self-hosting economically unattractive right now but signal that local, efficient runs are imminent. The broader claim is that distilled models from China will undercut frontier labs and change industry economics, while big tech’s willingness to overpay leaves them vulnerable to democratized, sustainable, frontier-capable alternatives.

Read on dgt.is320 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.