Anthropic
Claude 3.5 Sonnet
Announced Jun 20, 2024
A stronger everyday coding assistant
What could it do?
Help with code, writing and visual analysis, including charts and text in images.
What changed?
Anthropic reported improved capability and faster responses than Claude 3 Opus.
Coding assistance
Artifacts arrived alongside the model, giving generated documents, code and designs a dedicated workspace.
Where it fell short
Anthropic’s internal coding evaluation is not interchangeable with SWE bench Verified.
Original June 2024 release, not the October update. Google Cloud records June 20 availability; Anthropic’s page currently displays June 21. Check sources
- Version covered
- Original June release
- Context at launch
- 200,000 tokens
- Companion feature
- Artifacts preview
What this meant in practice
Separate the model from the workspace around it. Artifacts made output easier to inspect and revise, while the model generated that output. Both can make a workflow more useful, but they are different changes.
Draft a booking page
Describe the fields and layout, then inspect the resulting interface. Check keyboard access, validation and actual form submission separately. A convincing screen is only one part of a working site.
Common question
Does this include computer use?
This page covers June 2024. Do not assign features from later releases to this snapshot.
THE USEFUL CONTEXT
Why the original Claude 3.5 Sonnet mattered
Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.
The model and the workspace around it
Anthropic introduced Artifacts alongside Claude 3.5 Sonnet: generated documents, code and designs could be displayed beside the conversation. This is an important distinction for the timeline. The model produces material; the product determines how conveniently a person can inspect and revise it.
Imagine two assistants generating the same page. One returns a long block of code; the other lets you inspect a preview and request changes. The second experience might save time without proving that its underlying model has a higher score on every test. Comparing products means evaluating that experience as well as the model’s output.
A website example with real acceptance criteria
For an illustrative project, request a booking form for a fictional photography business. Specify a name field, email, preferred date and confirmation message. Start by inspecting the page on a narrow screen. Then use only the keyboard, submit empty fields and try a malformed email address. A polished screenshot does not answer any of those questions.
Next, decide what “submitted” means. Does the form only display a message, or does it save a request somewhere? This difference is easy to miss in an impressive demo. Before trusting generated software, define the behavior a person should experience and verify that behavior. These are proposed checks, not results from our own Claude test.
Keep the release version attached to the claim
This entry covers the original June 2024 release. If a comparison uses a later model with a similar name, it needs its own record. Otherwise a historical chart can accidentally make early versions look more capable by giving them features or scores that arrived months later.
For coding results, record the repository, the problem, the tools and the number of attempts. A model that suggests a correct edit in a chat window is being tested differently from an agent that searches files, runs tests and retries. Both can be useful, but the result describes the whole evaluation setup. The benchmark name alone is not enough to establish comparability.
FOLLOW THE EVIDENCE
What to watch next
Changes that would make this story worth revisiting:
- Releases that change how people inspect, edit or test generated work.
- Coding evaluations with reproducible tasks and clearly identified model versions.
Questions about this milestone
Does a working preview mean a website is ready to launch?
No. A preview demonstrates the visible interface. Saving data, handling errors, accessibility and behavior across devices require their own checks.
Why does this page discuss Artifacts separately?
A workspace feature can improve a user’s workflow without being a property of the model itself. Separating the two makes the history easier to understand.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models. The announcement page currently displays June 21, 2024.
Anthropic: Introducing Claude 3.5 Sonnet Google Cloud: June 20 launch availabilityBenchmark results
A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.
Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.
Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026