What happened?
In July 2025, Google DeepMind announced that an advanced version of its Gemini Deep Think model scored 35 of 42 points at the International Mathematical Olympiad (IMO), a gold medal level score DeepMind blog ↗. It solved five of the six problems perfectly, writing its proofs in ordinary language within the 4.5 hour contest time limit DeepMind blog ↗. IMO coordinators graded the answers using the same criteria as for student solutions DeepMind blog ↗.
The IMO is the world's top maths competition for school students. In 2025, 641 students from 112 countries took part, around 10% reached gold level, and five students scored a perfect 42 AFP report ↗. IMO President Gregor Dolinar said: "We can confirm that Google DeepMind has reached the much-desired milestone" DeepMind blog ↗.
This was a big step from the year before. In 2024, DeepMind's AlphaProof and AlphaGeometry 2 scored 28 points, but needed humans to translate the problems into a formal computer language and took days of computing DeepMind blog ↗. OpenAI also announced a 35 point result for its own experimental model, but that was graded by three former medalists it chose, not by IMO coordinators AFP report ↗ TechRepublic ↗.
A general purpose language model, not a special maths tool, produced complete proofs that official graders accepted under contest conditions.
Leapscope interpretation of the reported result.How did AI help?
The model read the official problem statements and wrote the proofs itself, end to end in natural language DeepMind blog ↗. DeepMind says it used parallel thinking, meaning it explored several possible solutions at once before settling on one, and was trained with new reinforcement learning techniques, a form of learning by trial and reward DeepMind blog ↗. Humans prepared its setup: it was given a curated collection of high quality maths solutions and general tips on approaching IMO problems DeepMind blog ↗.
Figures from the Google DeepMind announcement DeepMind blog ↗, confirmed in AFP reporting AFP report ↗.
The IMO checked the answers, not the system. DeepMind's own post says the review confirmed the answers were complete and correct but did not validate its system, processes or model DeepMind blog ↗. News reports added that organizers could not verify how much computing power the models used or whether humans were involved AFP report ↗. Also, the gold level model is not the one most people can use: the public version reaches bronze level on the same benchmark PYMNTS ↗.
Which fields could this affect?
The immediate value is in maths and AI research; wider uses are possible future value, and these connections are our assessment.
AI research
It is a clear, independently graded test of reasoning in a general language model DeepMind blog ↗. It also set a precedent for AI labs submitting results to official graders.
Explore scienceMaths learning
A version of Deep Think is available to Google AI Ultra subscribers in the Gemini app, though it is a faster, bronze level variant PYMNTS ↗. Students and teachers can test its explanations, with care.
Mathematical research
DeepMind gave the gold level version to a small group of mathematicians for feedback PYMNTS ↗. Whether it helps with real research questions has not yet been reported in these sources.
Explore scienceSolving open problems
Contest problems have known solutions written by experts. Doing well on them does not show an ability to solve questions nobody has answered.
What has been checked?
The evidence is official competition grading by IMO coordinators of the submitted answers, plus independent news reporting. There is no peer reviewed paper in these sources. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- IMO coordinators graded and certified the answers as worth 35 of 42 points DeepMind blog ↗.
- The proofs were written in natural language from the official problem statements DeepMind blog ↗.
- One of the six problems was not solved DeepMind blog ↗.
Still unknown
- How much computing power was used; the IMO could not verify it AFP report ↗.
- Whether the system itself works as DeepMind describes, since the IMO only checked the answers DeepMind blog ↗.
- How the gold level version performs on research mathematics, rather than contest problems PYMNTS ↗.
Evidence status: Official competition grading. Stage: Verified. Graded under official competition rules.
From contest gold to real research
This is our suggested way to follow the work, not a promised timetable.
Can I use it today?
Partly. Google AI Ultra subscribers can use a version of Deep Think in the Gemini app, but Google says it reaches bronze, not gold, level on the 2025 IMO problems PYMNTS ↗. The gold level version went only to a small group of mathematicians and academics PYMNTS ↗.
A few things you might be wondering
Did the AI win a real gold medal?
No medal was given. The IMO coordinators graded its answers as gold medal standard, 35 of 42 points DeepMind blog ↗.
Was it the only AI to score gold level?
No. OpenAI also reported 35 points, but its answers were graded by former medalists it chose rather than by IMO coordinators AFP report ↗ TechRepublic ↗.
Can I use the exact same model?
Not the gold level one. The public Deep Think in the Gemini app is a faster version that reaches bronze level PYMNTS ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01The developer's announcement with the score, grading details, IMO President quote and method.
News report on human results and the IMO's caveats about AI entries.
News report comparing the Google and OpenAI results and how each was graded.
Report on the public bronze level version and the gold level version given to mathematicians.