Early preview
A

Anthropic

Claude Haiku 4.5

Announced Oct 15, 2025

Capable coding in a smaller model

01

What could it do?

Anthropic reported coding performance similar to Sonnet 4 at lower cost and latency.

02

What changed?

Brought a newer capability level to the Haiku tier.

WHY IT MATTERED

Made frequent, smaller AI tasks more economical.

03

Where it fell short

Performance depends on the task; similar results on tests do not mean identical models.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

Developer announcement or release log

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
142.4index points
Tested variant: Claude Haiku 4.5Source interval: 139.4 to 144.1Variant date in source: Oct 15, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
17.0index points
Tested: Claude 4.5 HaikuIntelligence Index v4.3.2

Selected highest effort variant where available in the current table. Starred partial results are excluded. Claude fallback configurations can use other models when safeguards intervene.

Source: Artificial Analysis Intelligence Index ↗Download selected resultsChecked Oct 8, 2026
PUBLISHED BENCHMARK RESULT
1414rating points
Tested: claude-haiku-4-5-20251001Text Arena Overall · October 8, 2026Reported interval: 1412 to 1416146,351 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
PUBLISHED BENCHMARK RESULT
66.6%% resolved
Tested: Claude 4.5 Haiku (high)500 tasks · mini-SWE-agent v2.0.0Run: Feb 17, 2026 · Effort: highPublished submission; not marked as checked by the benchmark team

Only mini-SWE-agent v2.0.0 submissions, one run per model. Reasoning effort varies and is labelled. These are published submissions, not a rerun by Leapscope. Latest included release: February 5, 2026; latest run: February 26, 2026.

Source: SWE bench Verified ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.