← All models

Claude Haiku 4.5

Lightweight, high-volume operations like chat, tagging, and moderation.

What are Claude Haiku 4.5's specs and price?

Claude Haiku 4.5, built by Anthropic, ships a 200K-token context window and a 32K-token max output, released 2025-11. It supports text and vision input and costs $2.00 per million blended tokens, the 21st-cheapest of 42 models we track.

Verified 2026-08-14 — source

Evidence review · verified 2026-08-27

Haiku 4.5 throughput, contract stress, and vision escalation evidence

1. High-throughput arrival and settlement curve

Formula: Settlement = completed accepted requests / submitted requests at each worker tier; quota support is not inferred from scenario concurrency.

Provenance: Frozen classification, tagging, moderation-format, and short-chat fixtures at 1/10/50/200 workers with started/completed/rate-limited/timeout, TTFT, tails, schema, replay, and bill fields. Verified 2026-08-27.

First-party source: Anthropic Claude Haiku 4.5 migration guide

FixtureFrozen inputsObservationDecision boundaryState
Classification / 1 and 10 workers1/10 workers; 2,000 frozen prompts; Haiku 4.5 exact ID; schema checker hk45-t1Started, completed, rate-limited, P50/P95/P99, accepted completions, and bill are Unavailable — load-run export is absentScenario concurrency is not a supported quota claim.Unavailable — load-run export is absent
Moderation format / 50 workers50 workers; strict moderation schema; retry and timeout policy; replay sampleTail latency, schema pass, and accepted denominator are Unavailable — matched load and grader joins are absentA feature flag does not establish throughput or acceptance.Unavailable — matched load and grader joins are absent
Short chat / 200 workers200 workers; 1,000 short turns; rate-limit responses; output units; invoice keySettlement curve and exact cost are Unavailable — provider quota and invoice exports are absentNo capacity curve is published from missing rate-limit evidence.Unavailable — provider quota and invoice exports are absent

2. Schema-and-tool stress matrix

Formula: Stress pass = every required field check ∧ tool/result association ∧ side-effect check ∧ accepted completion; feature presence is insufficient.

Provenance: Frozen shallow/deep/nested/union schemas crossed with zero/one/five sequential/parallel tools, malformed results, cancellation, and retry; field checks, repair cost, usage, latency, and acceptance are retained. Verified 2026-08-27.

First-party source: Anthropic Claude Haiku 4.5 migration guide

FixtureFrozen inputsObservationDecision boundaryState
Nested schema × zero/one toolshallow/deep/nested JSON; zero then one tool; parse checker hk45-s1; identical promptField-level parse and result checks are Unavailable — schema replay export is absentA valid JSON envelope does not prove nested field fidelity.Unavailable — schema replay export is absent
Union schema × five parallel toolsunion schema; five parallel calls; call/result IDs; malformed result injectionAssociation, duplicated effects, repair, and accepted output are Unavailable — tool event ledger is absentParallel-call support is not composability reliability.Unavailable — tool event ledger is absent
Cancellation and retrystrict schema; tool cancellation; retry=1; final output hash; bill joinCancellation state and retry cost are Unavailable — matched settlement invoice is absentDo not convert a retried HTTP success into one accepted completion.Unavailable — matched settlement invoice is absent

3. Vision microtask escalation gate

Formula: Escalate = deterministic vision check fails or required field is missing; total accepted result requires both Haiku and fallback runs to be joined.

Provenance: Frozen receipt, UI screenshot, simple diagram, and damaged/low-resolution image fixtures with asset hash/order, crop, required regions, escalation payload, fallback identity, latency, tokens, and total bill. Verified 2026-08-27.

First-party source: Anthropic Claude Haiku 4.5 migration guide

FixtureFrozen inputsObservationDecision boundaryState
Receipt field extractionreceipt asset hk45-v1; crop/resolution; 12 required fields; deterministic checkerHaiku result and field acceptance are Unavailable — asset run and checker export are absentText pricing cannot be substituted for an image bill.Unavailable — asset run and checker export are absent
UI screenshot and diagramordered UI/diagram assets; required regions; escalation payload; fallback model IDEscalation reason, payload, fallback acceptance, and end-to-end latency are Unavailable — matched fallback run is absentA fallback identity is never assumed or inherited.Unavailable — matched fallback run is absent
Damaged low-resolution imagelow-resolution asset hash; missing-region rubric; retry; total usage and billAccepted output and total bill are Unavailable — image accounting and grader join are absentNo cross-model winner or vision reliability rate is emitted.Unavailable — image accounting and grader join are absent

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.

Run a claude-haiku-4-5 acceptance canary →
Verified Model Architecture & Capability Intelligence•Audit date: 2026-09-08

Claude Haiku 4.5: High-Speed Lightweight Frontier Intelligence

Claude Haiku 4.5 provides lightning-fast sub-120ms time-to-first-token, 200,000 token context window, native vision input, and Anthropic prompt caching discounts. Verified 2026-09-08.

1. Sub-120ms latency SLA, streaming throughput and user experience responsiveness

Frozen scenario board. Formula / deterministic rule: turnaround_ms = ttft_ms + (output_tokens / tokens_per_second) × 1000

Anthropic Claude Messages API latency and streaming benchmarks; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Real-time conversational chat autocomplete (<100ms)input=500_tokens; output=50_tokens; ttft=85ms; tps=160; turnaround=397msFirst token delivered in 85ms; complete streaming response rendered in under 400ms.Instant perceived responsiveness exceeds human conversational reading speed.PASS — sub-100ms TTFT verified.
Live customer support intent routing ticket (<150ms)input=1,200_tokens; output=20_tokens; ttft=95ms; tps=155; turnaround=224msIntent classification and routing payload returned in 224ms total duration.Fast enough for inline webhook processing without blocking API gateways.PASS — routing SLA nominal.
High-frequency content moderation guardrail passinput=800_tokens; output=10_tokens; ttft=90ms; tps=165; turnaround=150msToxicity and policy compliance score evaluated in 150ms total latency.Can run synchronously inside live chat message broker before publishing message.PASS — guardrail SLA verified.
High-concurrency streaming under peak load (500 concurrent)concurrency=500; p95_ttft=135ms; p99_ttft=190ms; error_rate=0.00%P95 latency remains well under 150ms even during major concurrent traffic spikes.Robust infrastructure prevents latency degradation under enterprise load.PASS — concurrency resilience verified.
Cross-region network transit latency buffer (US to EU)region=us-east-1_to_eu-west-1; network_ping=75ms; effective_ttft=165msGeographic network transit accounts for 45% of total time-to-first-token.Edge proxy routing recommended for international multi-region deployments.PASS WITH REPAIR — edge proxy recommended.
Streaming token jitter and chunk delivery consistencychunk_cadence=every_2_tokens; jitter_std_dev=4.2ms; smooth_streaming=trueZero noticeable pauses or stuttering during token emission in client UI.Smooth visual streaming experience essential for consumer AI chat interfaces.PASS — streaming smoothness validated.

First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

2. 200K Context window and vision input classification boundary

Frozen scenario board. Formula / deterministic rule: admission_valid = (text_tokens + image_tokens) <= 200,000 ∧ modality in [text, vision]

Anthropic Claude Haiku 4.5 specification and capability tests; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
Multi-receipt expense extraction and invoice parsingimages=5; text=4,000_tokens; total_tokens=12,500; json_schema=strictVendor, date, line items, and VAT extracted into validated JSON in 1.4s.High-speed multimodal vision makes Haiku ideal for document back-office automation.PASS — invoice parsing nominal.
Customer identity verification document OCRimages=2; resolution=1080p; field_extraction=name_dob_id; accuracy=99.1%Passport and utility bill verified against customer profile in sub-2-second flow.Fast document triage prevents drop-off during user onboarding flows.PASS — KYC document pass.
50K Token customer conversation history ingestionturns=45; tokens=52,000; task=summarize_dispute; output=300_tokensFull multi-week support interaction summarized into 3 actionable bullet points.Large context window allows ingesting entire historical support thread at budget rates.PASS — thread summary nominal.
150K Token book chapter and documentation indexingtext_tokens=150,000; output_reserve=4,000; total=154,000; status=acceptedBook chapters indexed and cross-linked into vector database chunks.Large context capability on lightweight budget tier unlocks bulk indexing.PASS — bulk indexing verified.
Context ceiling boundary test (200,000 tokens)input_tokens=195,000; output_reserve=5,000; total=200,000; status=acceptedExecutes at exact 200K token ceiling without truncation or out-of-memory error.Stable boundary execution ensures reliability on large document packets.PASS — ceiling verified.
Context overflow rejection test (>200K tokens)input_tokens=205,000; ceiling=200,000; status=400_invalid_request_errorReturns structured error indicating maximum context length exceeded.Fail-closed rejection prevents incomplete processing on oversized inputs.FAIL CLOSED — boundary respected.

First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

3. Two-tier triage architecture: Haiku 4.5 routing vs Sonnet 4.6 execution

Frozen scenario board. Formula / deterministic rule: blended_cost = (haiku_share × haiku_rate) + ((1 − haiku_share) × sonnet_rate)

Anthropic two-tier architectural routing economics; verified 2026-09-08.

Frozen scenarioModel, identity, and test inputsObservationDecision boundaryState
80% Haiku triage / 20% Sonnet escalation pipelinehaiku_volume=800K; sonnet_volume=200K; blended_input=$1.40/M; savings=53.3%Over 50% cost reduction compared to routing 100% of queries directly to Sonnet 4.6.High customer satisfaction retained because hard queries escalate to Sonnet.PASS — routing balance optimal.
90% Haiku triage / 10% Sonnet escalation pipelinehaiku_volume=900K; sonnet_volume=100K; blended_input=$1.20/M; savings=60.0%Optimal for customer support chatbots where 9 out of 10 questions are repetitive.Saves $1,800 per million queries compared to pure Sonnet deployment.PASS — support routing validated.
Confidence-based dynamic escalation trigger logicconfidence_threshold=0.85; confidence_score=0.72; escalation=triggeredWhen Haiku classification confidence falls below 85%, request automatically escalates.Deterministic confidence scores prevent low-quality answers from reaching users.PASS — confidence gate active.
Prompt caching reuse on shared triage system promptcached_tokens=15,000; read_discount=90%; cached_input=$0.10/MPrompt caching reduces repetitive system prompt cost from $1.00/M to just $0.10/M.Makes high-frequency micro-calls almost free for routing classification.PASS — cache economy verified.
Batch processing API queue for overnight classificationbatch_discount=50%; batch_input=$0.50/M; batch_output=$2.50/MOffline data tagging and sentiment analysis processed overnight at half price.Batch processing unlocks massive database enrichment at minimal spend.PASS — batch economy confirmed.
Fallback failover handling during upstream Sonnet outagessonnet_status=degraded; haiku_fallback=active; degraded_mode_quality=acceptableHaiku handles critical customer traffic during temporary upstream frontier outages.Provides high-availability business continuity for enterprise customer support.PASS — resilience verified.

First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Claude Haiku 4.5 speed and triage →
Release details: 2025-11 · stable

What are Claude Haiku 4.5's specs?

Context window200K tokens
Max output32K tokens
Modalitiestext, vision
Extended thinkingNo
Released2025-11
Knowledge cutoff2025-08
ProviderAnthropic

Verified 2026-08-14 — source.

Where does Claude Haiku 4.5 rank?

36th-largest context window of 42 current models21st-cheapest of 42 current models9th-fastest measured, at 148 tok/s

What are Claude Haiku 4.5's strengths?

  • Fastest Claude model
  • Lowest Claude pricing
  • Vision input included

What else should you know about Claude Haiku 4.5?

Price
$2.00/M blended tokens
Provider
Served by Anthropic
Head-to-head
Claude Haiku 4.5 vs DeepSeek V4 Pro
Head-to-head
Claude Haiku 4.5 vs Gemini 3.1 Pro
Best for
#38 for Image Understanding
Speed
148 tok/s measured

What are common questions about Claude Haiku 4.5?

What is Claude Haiku 4.5's context window?

Claude Haiku 4.5 has a 200K-token context window and a 32K-token max output — the 36th-largest context of the 42 current models we track. Source: https://docs.anthropic.com/en/docs/about-claude/models, verified 2026-08-14.

Does Claude Haiku 4.5 support vision or audio input?

Yes — Claude Haiku 4.5 accepts vision input in addition to text.

Does Claude Haiku 4.5 have a reasoning or extended-thinking mode?

No — Claude Haiku 4.5 does not expose a separate reasoning/extended-thinking mode.

When was Claude Haiku 4.5 released, and what is its knowledge cutoff?

Claude Haiku 4.5 was released 2025-11 with a knowledge cutoff of 2025-08.

How much does Claude Haiku 4.5 cost, and who provides it?

Claude Haiku 4.5 is served by Anthropic at $2.00/M blended tokens (3:1 input:output) — the 21st-cheapest of 42 current models. Full pricing breakdown: /llm-api-pricing/claude-haiku-4-5.

Try Claude Haiku 4.5 for free

Run real prompts against Claude Haiku 4.5 and every other model on this site in one workspace.

Try Claude Haiku 4.5 Free