hn.today

Which AI model is best for coding? A hands-on comparison

flaviocopes.com3 points0 comments
Screenshot of Which AI model is best for coding? A hands-on comparison

A practical comparison of coding-focused AI models in September 2026 that tests how they perform on real coding tasks and what that performance costs. The clear winner for most agent-driven coding work is Claude Opus 5.5: it tops Cursor’s coding leaderboard, is faster in the author’s bake-off, and is cost-effective at the medium effort setting (max effort ballooned costs dramatically). GPT-5.6 Sol is a close second for precise, well-specified edits and used the fewest tokens in tests; GPT-6 (Sol, Luna, Astra) is new and promising but hadn’t been widely available in the author’s tools at test time. Grok 4.7 is cheaper per-token but, because of billing quirks, can cost more per task; open-weight options (GLM-5.3, Kimi K3 hosted, Muse Glimmer, Gemma 4 locally) are viable for privacy or self-hosting but don’t match closed models on long agent runs.

The bake-off method matters: a tiny JavaScript project with two bugs and a missing feature was given to six models via the same Cursor CLI agent, then checked with hidden tests the models never saw; each model ran three times. Every model passed the explicit spec, but differences showed up in edge cases, consistency between runs, token use, runtime and total task cost. The write-up includes per-model price/context limits, effort settings, leaderboard context (CursorBench 4.0) and practical recommendations: Opus 5.5 for most agent work, Fable 5.1 for long tasks or front-end, GPT-5.6/GPT-6 Sol for deadline-sensitive precise edits, Luna tiers for high-volume jobs, and open weights for offline/private use.

Read on flaviocopes.com0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.