Claude Haiku 4.5

Anthropic released 2025-10-15 proprietary

$2.00 per million tokens, balanced

Gemini 3.7 Flash is both better and cheaper.

It scores +26.1 higher and costs 63% less ($0.750/M against $2.00/M) on this workload — and it does everything this model does.

DeepSeek V4 Flash 0731 is cheaper still (95% less) but drops no image, file input.

The same model via batch is 50% cheaper ($1.00/M) — same weights, different latency.

Our take

editorial — not a measurement

The cheap Anthropic tier at $1/$5, and the odd one out in the current lineup — it still uses extended thinking rather than adaptive thinking, and it is capped at a 200k context and 64k output while every 4.6-and-later model gets 1M. That makes it a classification and routing model, not a small general assistant. At $1 input it is roughly 20x the price of the cheapest capable open-weight models, so it only makes sense when you specifically want Anthropic's safety behaviour and refusal profile at low cost.

Strengths

  • $1/$5 with Anthropic's alignment and tool-use behaviour
  • Fastest Claude latency profile
  • Full caching and 50% batch discount support

Weaknesses

  • 200k context and 64k output — a fifth of the current generation
  • Extended thinking rather than adaptive thinking
  • Far more expensive than equivalent open-weight small models
  • Older knowledge cutoff (training data to Jul 2025)

Reach for it when

  • Classification and extraction
  • Router / triage models in front of a larger model
  • High-volume support ticket handling

Avoid it if

  • You need long context
  • Price per token is the deciding factor

Sources: platform.claude.complatform.claude.com

Every price dimension

Input$1.00/M
Output$5.00/M
Cached input$0.100/M
Cache write$1.25/M
Cache write 1h$2.00/M
Web search$0.01/call
Batch discount50%

Independent scores

Intelligence29.9
Coding43.9
Agentic16.5

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window200K
Max output64K
Input modestext, image, file
Tool useyes
Reasoningoptional
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

Markdown for LLMs