hn.today

Claude Haiku 5.5

platform.claude.com13 points1 comments
Screenshot of Claude Haiku 5.5

Claude Haiku 5.5 is a lightweight, latency-optimized large language model built for high-volume tasks like classification, extraction, routing, and subagent work. It offers a 1 million token context window, up to 128k output tokens (300k in Batch API beta), and adaptive “thinking” with an effort parameter to control depth. The model uses a newer tokenizer (roughly 30% higher token counts compared with Haiku 4.5), defaults to medium effort, and is positioned as the fastest and most cost-effective option among the 5.5 family. Reliable knowledge and training data cutoff are June 2026, it launched October 7, 2026, and it is available across Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. Thinking blocks are account-bound or limited to linked accounts, and non-default temperature/top_p/top_k values are not allowed (they cause a 400 error).

Pricing is tiered by token volume: input starts at $0.10 per million tokens up to 100k and $0.50/MTok above that; output is $0.50/MTok up to 100k and $2.50/MTok above. Cache write/read rates and 5m/1h cache write options are billed separately, and Batch API usage receives a 50% discount on input and output. Documentation highlights migration guidance, prompting best practices, latency reduction techniques, and platform limits; it also lists model IDs, versioning, deprecation timelines, and safety/system prompt references.

Read on platform.claude.com1 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.