Anthropic has released Claude Sonnet 5.5, which scores 56 on the Artificial Analysis Intelligence Index - two points behind Opus 5.5 (max) and 18 points above Sonnet 5 - putting it at #2 on the leaderboard. Pricing remains unchanged from Sonnet 5 at $2/$10 per million input/output tokens (with cache write/read charges), but Sonnet 5.5’s higher output volumes raise its average Cost per Task to about $7.60, roughly 50% above Sonnet 5. The model retains a 1‑million‑token context window and offers five effort settings; fallback to Sonnet 5 occurs in about 0.1% of tasks. A pre-release bug that affected structured-output requests has been fixed for the public release.
Substantively, Sonnet 5.5 matches leading models on agentic terminal use and knowledge-work benchmarks - Terminal‑Bench 4.0 reaches 64%, and it hits parity with Opus 5.5 on AA‑Briefcase, GDPval‑AA and AutomationBench‑AA - but it achieves those results using far more output tokens. At max effort it produces ~193k output tokens per Intelligence Index task, the highest measured and ~60% above Opus 5.5 (max) or ~7x GPT‑6 Astra (max). It trails Opus 5.5 on factual accuracy and scientific reasoning (AA‑Omniscience 54% vs 66%, lower scores on Humanity’s Last Exam and SciCode) and sits slightly off the Intelligence vs. Cost Pareto frontier, where GPT‑6 Sol and other configs deliver similar performance at lower token cost.
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.