Early preview
A

Anthropic

Claude 3.7 Sonnet

Announced Feb 24, 2025

Reasoning and coding in one model

01

What could it do?

Produce a standard response or use extended thinking for a harder task.

02

What changed?

Users could choose a reasoning mode within the same model; API users could also set a thinking budget.

WHY IT MATTERED

Hybrid reasoning

Claude Code introduced a terminal workflow for reading files, editing code and running tools.

03

Where it fell short

Coding benchmark results depend on the tools, retry budget and evaluation setup around the model.

Which release does this page cover?

February 2025 release, alongside the limited Claude Code research preview. Check sources

Reasoning modes
Standard and extended
Companion tool
Claude Code
Tool access at launch
Limited research preview

What this meant in practice

A model and a coding agent are different layers. The model proposes actions; the surrounding tool lets it inspect files or execute tests. More thinking and more tool access also change the time and cost of a task.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Investigate a failing test

Give an agent a repository and a specific failure. Ask it to explain the cause, propose a change and run the relevant checks. Review the change before treating the task as complete.

Common question

Is Claude Code the model itself?

No. Claude Code is a tool that uses a model to work in a software environment.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

Anthropic: Claude 3.7 Sonnet and Claude Code

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
141.2index points
Tested variant: Claude 3.7 SonnetSource interval: 138.8 to 142.7Variant date in source: Feb 24, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1388rating points
Tested: claude-3-7-sonnet-20250219-thinking-32kText Arena Overall · October 8, 2026Reported interval: 1384 to 139238,162 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
ONE EMAIL A WEEK

New AI tools worth paying for.

Price changes, new free plans and the picks we would change. Short, honest, no sponsors in the picks.