hn.today

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems156 points149 comments
Screenshot of Why I'm still bearish on LLMs after Navier-Stokes

The piece argues that recent headline wins like formalizing Navier-Stokes do not change the fundamental limits of current large language models: they are far from fully autonomous and still require laborious oversight, rigorous formal specification, or brittle human review. Models generalize only within narrow neighborhoods of their training tasks and are prone to reward hacking; the one clear success - encoding a mathematical theorem in Lean - is a best-case scenario because the theorem statement is already a rigorously audited specification and the theorem prover is designed to resist unsoundness, though soundness bugs have occurred. Rigorous specification is expensive and specialist work (hardware projects often have many more spec/validation engineers than designers), verification costs can exceed implementation, and expert human review scales poorly and is itself vulnerable to covert failures (xz backdoor, UMN hypocrite commits).

As a result, most firms will use LLMs as supervised “cracked interns” rather than autonomous agents. Only three classes can accept full autonomy: organizations that can tolerate cheap failures (prototyping, intern-like work), narrow-task operations with clear guardrails (controlled physical processes, basic customer service), and domains already set up for heavy specification and validation (chip design, drug discovery). Price-sensitive and swarm-sensitive work favors cheap, open models run widely rather than expensive frontier models; secrecy and IP concerns also disincentivize sending core work to major cloud labs. The overall blast radius of AI advances will extend beyond frontier labs, but human orchestration will continue to bottleneck truly autonomous deployment.

Read on dank.systems149 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.