hn.today

DeepSeek is training a 2T-parameter model and plans to build an 8T-parameter one

twitter.com12 points0 comments
Screenshot of DeepSeek is training a 2T-parameter model and plans to build an 8T-parameter one

DeepSeek is shifting its AI training strategy to prioritize domestically produced accelerators, with CEO Liang Wenfeng telling investors that moving away from foreign chips is now a major priority. The company expects Huawei to begin delivering relevant hardware as early as Q4, and is actively training a 2 trillion-parameter model now with an explicit roadmap to scale up to an 8 trillion-parameter model in the future. Those parameter targets and the domestic-first procurement stance signal a sizable change in both model ambition and the underlying infrastructure choices.

The announcement frames the decision as primarily a supply-chain and capital expenditure problem: sourcing domestically reduces exposure to export controls but forces attention to wafer, interconnect and memory lead times. Observers point out that memory and other components still need to come from somewhere, so the shift will affect procurement, timelines and cost structure as much as model design. For investors and builders, the news matters because it ties model scale directly to hardware availability, delivery schedules (Huawei Q4), and the broader geopolitics of chip supply.

Read on twitter.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

Xiaomi MiMo v2.6

Xiaomi MiMo v2.6

Xiaomi MiMo v2.6 is an open-source language model that demonstrates diverse capabilities, including use in scientific environments and multimedia tasks. The release is seen as a significant step forward in lightweight models at an affordable cost. (mimo.xiaomi.com)

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.