THE MEASUREMENTS
A score needs a little context.
Three different tests answer three different questions.
Benchmarks are the main way AI labs and independent testers compare models, but each one measures something different. Arena ranks models by human preference in blind head to head votes. SWE bench Verified checks whether a model can fix real software issues. The Artificial Analysis Intelligence Index combines several evaluations into one number. These guides explain what each score means, where it comes from and why combining them needs care.
Artificial Analysis
It combines multiple evaluations into an overall index. The benchmark mix, weights and grading methods can change between versions.
Explore BENCHMARK GUIDEArena
Arena asks people to compare anonymous model responses. Those preferences contribute to relative model rankings.
Explore BENCHMARK GUIDESWE bench Verified
SWE bench tests whether a system can produce a patch for a real software issue. Verified is a selected set of 500 problems checked by engineers.
Explore