logoSeailife
  • Submit Project
  • Pricing
  • Blog
Sign inSign up
Sign in
logoSeailife

© 2026 Seailife. All rights reserved.

Built with Open Launch - The first complete open source alternative to Product Hunt.

Powered by Open-LaunchPowered by Open-Launch

Discover

  • Trending
  • Categories
  • Submit Project

Resources

  • Pricing
  • Sponsors
  • Blog

Legal

  • Terms of Service
  • Privacy Policy

Connect

  • [email protected]
Claude Haiku 5.5 Logo

Claude Haiku 5.5

Artificial IntelligenceDeveloper Tools
Visit
0 upvotes
Claude Haiku 5.5 - Product Image

Claude Haiku 5.5 is Anthropic's small model for high-volume work: classify, summarize, and run subagents, with two prices split at 100,000 tokens.

Two prices, split at 100,000 tokens

Anthropic published Claude Haiku 5.5 on 7 October 2026. The rate depends on how long the prompt is. A prompt of 100,000 tokens or less is one price. A longer prompt is five times the input price. On the previous Haiku, Anthropic says about 90% of requests were in the shorter band. The figures below are per million tokens, from that announcement and the Claude pricing page.

  • Prompt up to 100,000 tokens: input $0.10, output $0.50, cache read $0.01, 5-minute cache write $0.125, 1-hour cache write $0.20.
  • Prompt over 100,000 tokens: input $0.50, output $2.50, cache read $0.05, 5-minute cache write $0.625, 1-hour cache write $1.00.
  • Claude Haiku 4.5, for comparison, is one rate for any prompt length: input $1, output $5, cache read $0.10, 5-minute cache write $1.25.
  • The Batch API takes 50% off input and output. That discount applies on both sides of the 100,000-token line.

The sticker price for a short prompt is 90% below Haiku 4.5, and 50% below it once the prompt crosses 100,000 tokens. Anthropic's own average is lower than that sticker: around 75% less to run. The gap is the tokenizer. Haiku 5.5 uses the same newer tokenizer as Claude 4.7 and later, and the same text counts as about 30% more tokens than it did on Haiku 4.5. The exact increase depends on the text. The same announcement also cut Claude Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens, which they say makes most agentic Sonnet work about 20% cheaper. That cut is a Sonnet change, not a Haiku rate.

The id, the window, and where it runs

On the model overview the API id is claude-haiku-5-5. There is no date suffix and no alias on that page.

  • Claude API, Google Cloud, Microsoft Foundry, and Claude Platform on AWS: claude-haiku-5-5.
  • Amazon Bedrock: anthropic.claude-haiku-5-5.
  • Context window 1 million tokens, up from 200,000 on Haiku 4.5. Max output 128,000 tokens on the synchronous Messages API, up from 64,000. The Message Batches API can go to 300,000 output tokens with the beta header output-300k-2026-03-24.
  • Input is text and images. Output is text. Reliable knowledge cutoff is June 2026. Status is active, and the overview says retirement is not sooner than 7 October 2027.
  • Adaptive thinking is on by default. Default effort is medium. The overview calls the comparative latency the fastest in the current lineup at standard speed. The announcement's footnote says it still runs less quickly than Opus models in Fast Mode.

What breaks if you only change the model string

The what's-new page and the migration guide list changes that return an error, or that change the bill, if a Haiku 4.5 client is pointed at this id.

  • A non-default temperature, top_p, or top_k returns a 400. Omit them.
  • An assistant prefill returns an error. The messages array has to end on a user turn.
  • Manual extended thinking with budget_tokens returns an error. Thinking is adaptive. You can still turn it off with thinking: {"type": "disabled"} at high effort or below. Effort is the control they recommend for trading quality against speed and cost.
  • Thinking tokens count toward max_tokens. An old small limit can stop after a thinking block and before any text.
  • A response can start with a thinking block even when the request never mentions thinking. Read blocks by type, not by position. Thinking text is omitted unless thinking.display is summarized.
  • Stored thinking blocks work only in the account that produced them, or in an account linked to it. Replaying a saved conversation through a different account drops them.
  • On the Claude API and Google Cloud, computer use moves from computer_20250124 to computer_toolset_20260801. Browser use is a new tool on those two platforms.
  • Safety classifiers can decline a request with stop_reason: "refusal". The what's-new page says server-side fallback is not available, so the client has to handle that stop reason.

Scores Anthropic printed, not a retest

The announcement's benchmark table is Anthropic's own run. This page did not call the model. Where a cell was blank on their table, it is left blank here. Sonnet 5.5 is the column they label "for reference."

  • GDPval-AA v2.1: Haiku 5.5 1620, Haiku 4.5 735, GPT-6 Luna 1437, Sonnet 5.5 1840.
  • AA-Briefcase v1.1: 1578, 614, 1336, 1824.
  • OSWorld 2.1, offline subset: 72.4%, 15.7%, 48.9%, 83.9%.
  • Humanity's Last Exam, no tools: 45.9%, 10.2%, no Luna figure, 56.9%. With tools: 57.4%, 18.7%, no Luna figure, 64.5%.
  • Terminal-Bench 4.0: 39.2%, 0.0%, 16.4%, 70.6%.
  • FrontierCode 1.1 (Main): 46.4%, no Haiku 4.5 figure, 42.4%, Sonnet 5.5 52.1% at Xhigh.
  • Chartography, no tools: 46.4%, 6.4%, 29.1%, 61.6%.

On that same page they draw the job split in one sentence: Sonnet 5.5 and Opus 5.5 stay the better choice for complex agentic coding, the kind Terminal-Bench 4.0 measures. Haiku 5.5 is for narrower work that used to be too expensive to run often, such as compaction, summarization, and subagents. It is also their first Haiku with an adjustable effort setting, so the same id can be turned toward cost or toward the score.

Early tests the customers put their names on

These are early tests Anthropic printed on the announcement, with the company and the person attached. They are not an independent audit, and they are not a result from this site.

  • Asana, Aaron Vinh, Staff Software Engineer: on the AI Teammates suite (bug triage, project setup, searching a portfolio for high-risk or overdue work), more than 30% lower latency for task completion and up to 2.5 times faster inference per agent turn than the model they use today.
  • HubSpot, Ze’ev Klapow, Distinguished Software Engineer: 92.8% on their simulated CRM suite, averaged over three runs, the best score they report seeing on that suite. On a stale-record audit, it was the fastest of the models they tested, with the highest hit rate and the lowest false-positive rate.
  • AlphaSense, Daniel Campos, Distinguished Engineer: Ask in Document does about 8 million calls a week. On 400 queries, Haiku 5.5 scored 0.84 against 0.76 for Haiku 4.5, which they call a statistically significant improvement.
  • Box, Yashodha Bhavnani, VP of AI Products: 11 points higher than Haiku 4.5 at about half the latency, in early testing aimed at cost reports, financial summaries, and weekly reviews.
  • Rogo, Alex Wang, Applied AI: they describe the fit as short, high-volume work, lookups, subagents, and summaries. Their example is a Haiku subagent pulling a segment-revenue line out of a 10-K while a larger model builds the deck.
  • Cognition, Walden Yan, Co-Founder and CPO: with Haiku 5.5 as the sidekick in Devin Fusion, they report a FrontierCode score of 66.2 while cutting cost and latency, and say it can be tried in the Devin CLI with Opus 5.5 as the lead.

The safety line they drew

The announcement says alignment evaluations improved against Haiku 4.5, with fewer instances of misaligned behavior and a lower willingness to cooperate with misuse. The detail is in the Haiku 5.5 system card, which also says the knowledge cutoff is June 2026. Cybersecurity safeguards are more restrictive than Haiku 4.5 and somewhat less restrictive than the ones on other recent models: they allow a wider set of defensive tasks than the Sonnet 5.5 safeguards, and they still block penetration testing. Biology safeguards match Sonnet 5, Sonnet 5.5, and Opus 5: research questions are allowed, and requests judged likely to cause harm are restricted. Wider biology or cyber work goes through the Life Sciences Verification Program or the Cyber Verification Program.

The same post says a monthly API credit rolls out this week for Claude Max and Team subscribers, for experiments on the Claude Platform. Max 5x gets $100 a month, Max 20x gets $200, and Team gets up to $500 pooled across users. The credit works on any Claude model, not only Haiku. Anthropic's note is the Help Center article linked from the announcement.

Comments

Publisher

Seailife

Seailife

Launch Date
2026-10-08
Platform
api
Pricing
paid
Socials

Sponsors

Become a Sponsor

Get your brand featured here