Early preview
G

Google

Gemini 2.5 Pro

Announced Mar 25, 2025

Reasoning across long and visual inputs

01

What could it do?

Use a thinking approach for reasoning and coding while processing text, images, audio and video.

02

What changed?

Google combined a stronger base model and further training with built in thinking capabilities.

WHY IT MATTERED

Multimodal reasoning

The initial experimental release brought reasoning and a one million token context window into the same model.

03

Where it fell short

Results from an experimental version should not be silently assigned to a later stable version.

Which release does this page cover?

The March 25, 2025 experimental release, not every later Gemini 2.5 Pro version. Check sources

Release stage
Experimental
Context at this release
1 million tokens
Output focus
Reasoning and code

What this meant in practice

“Thinking” describes a model’s approach, not a promise of correctness. For a useful comparison, record the exact version and the task conditions, then measure whether its answer solves the problem.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Inspect a large project

Ask which files are relevant to a bug before requesting a patch. Check the proposed file references and test the change. A large input window does not remove the need for a clear task.

Common question

Was the two million token window available at this launch?

The announcement described two million tokens as coming later; the initial release had one million.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

Google: Gemini 2.5 experimental launch

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
144.2index points
Tested variant: Gemini 2.5 Pro (Mar 2025)Source interval: 141.9 to 147.0Variant date in source: Mar 31, 2025

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1445rating points
Tested: gemini-2.5-proText Arena Overall · October 8, 2026Reported interval: 1443 to 1447125,379 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.