Anthropic
Claude Sonnet 4.5
Announced Sep 29, 2025
Longer coding and computer tasks
What could it do?
Anthropic reported gains in coding, computer use and complex agent workflows.
What changed?
Launched alongside its Agent SDK and updates to Claude Code.
Developers gained tools for building agents that work across multiple steps.
Where it fell short
Long task demonstrations do not establish reliability on every project.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.
Developer announcement or release logBenchmark results
A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.
Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.
Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026Only mini-SWE-agent v2.0.0 submissions, one run per model. Reasoning effort varies and is labelled. These are published submissions, not a rerun by Leapscope. Latest included release: February 5, 2026; latest run: February 26, 2026.
Source: SWE bench Verified ↗Download selected resultsChecked Oct 8, 2026