Early preview

THE MEASUREMENTS

A score needs a little context.

Three different tests answer three different questions.

Benchmarks are the main way AI labs and independent testers compare models, but each one measures something different. Arena ranks models by human preference in blind head to head votes. SWE bench Verified checks whether a model can fix real software issues. The Artificial Analysis Intelligence Index combines several evaluations into one number. These guides explain what each score means, where it comes from and why combining them needs care.

BENCHMARK GUIDE

Artificial Analysis

It combines multiple evaluations into an overall index. The benchmark mix, weights and grading methods can change between versions.

Explore
BENCHMARK GUIDE

Arena

Arena asks people to compare anonymous model responses. Those preferences contribute to relative model rankings.

Explore
BENCHMARK GUIDE

SWE bench Verified

SWE bench tests whether a system can produce a patch for a real software issue. Verified is a selected set of 500 problems checked by engineers.

Explore
The capability chart has sourced Epoch AI estimates. Artificial Analysis and Arena now include sourced 2026 results. SWE bench contains a smaller set of comparable runs, with its coverage cutoff shown in the chart.
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.