hn.today

2026 in LLMs (So Far)

simonw.substack.com20 points3 comments
Screenshot of 2026 in LLMs (So Far)

A brisk retrospective of 2026 LLM developments that stitches together model releases, product experiments, and cultural effects observed up to September. The keynote traces a timeline from late‑2025 releases (Claude Opus 4.5, GPT‑5.1) that pushed coding agents from “often wrong” to “reliable enough for daily use,” through a January burst of “AI mania” projects (a Python JavaScript interpreter and a WebAssembly runtime) that tested the limits of agentic tooling. A running, tongue‑in‑cheek benchmark - asking models to “generate an SVG of a pelican riding a bicycle” - is used to illustrate incremental improvements across models, and later mentions include Claude Opus 5.5, Gemini 3.1 Pro, and the headline topics of GPT‑6 Sol, GPT‑6 Luna and an emerging price war, signaling intense vendor competition.

Substantively, the writeup argues that coding agents reshaped software development practices: enormous fast‑moving projects like OpenClaw (100k+ commits) spawned a new category of “Claw” personal agents, drove hardware demand (Mac Minis as “aquaria”), and even produced ephemeral social experiments (MoltBook) that exploded then collapsed. Security, sandboxing and the question of human review dominated conferences - exemplified by StrongDM’s “Dark Factory” rules banning human code writing and review - while cultural effects include coined terminology (“Deep Blue”) for engineer ennui. The overall picture is rapid iteration, operational experimentation, and growing pains as agents move from novelty to infrastructure.

Read on simonw.substack.com3 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.

2026 in LLMs (So Far) · hn.today