AI × MATHEMATICSExplained simply

Ten unpublished math problems. One AI sprint.
Meet the First Proof attempts.

Mathematicians set ten fresh research problems to test whether AI can write real proofs. OpenAI shared its model's attempts. Here is what held up, what did not, and why the test matters.

WHERE THIS STANDS
  1. Claim
  2. Verified
  3. Usable
  4. In use

Proof attempts shared publicly; independent checking is not complete.

What moves it next: Moves to Verified when outside experts or independent teams confirm the result. How we decide

THE BREAKTHROUGHProof attempts on ten research problems
THE TEAMOpenAI, internal unreleased model
WHERE IT STANDSPartly checked; some attempts wrong
01 · THE BREAKTHROUGH

What happened?

In February 2026, OpenAI published its internal model's attempts at the ten First Proof problems, a challenge built to test whether AI can produce checkable proofs in specialised research fields OpenAI post ↗. Based on expert feedback, OpenAI said at least five attempts, for problems 4, 5, 6, 9 and 10, had a high chance of being correct OpenAI post ↗.

First Proof was created by eleven prominent mathematicians, including Fields Medal winner Martin Hairer, who each contributed a lemma from their own unpublished work Scientific American ↗. A lemma is a smaller theorem used as a step toward a bigger result. Because the answers were not online, an AI could not simply look them up. The problems were released on February 5, 2026, and the organisers released their own solutions on February 13 First Proof site ↗.

OpenAI shared its attempts on February 14 and explained its process in a post dated February 20 OpenAI post ↗. It first believed its answer to problem 2 was likely correct, but after official commentary and community analysis it now believes that answer is wrong OpenAI post ↗. Chief scientist Jakub Pachocki had earlier said six of ten had a high chance of being correct, and mathematicians pointed to possible gaps in at least one of those Scientific American ↗.

What are the three pieces of the story?

The test

Ten research level lemmas from different fields, set by working mathematicians who already knew the answers Scientific American ↗. Each needed a full written argument, not a short numerical answer OpenAI post ↗.

The attempts

OpenAI's unreleased model produced written proofs for all ten problems OpenAI post ↗. OpenAI says around five are likely correct, while others were still under review OpenAI post ↗.

The checking

Organisers checked AI proofs themselves, but outside submissions with human help were harder to judge Scientific American ↗. A later Nature report said the February trial results were not officially verified by the First Proof team Nature news ↗.

THE REASON TO BE EXCITED

The value here is less the score and more the format: fresh problems with known answers make it possible to check AI mathematics in public rather than take a company's word for it.

Leapscope interpretation of the reported result.
02 · AI’S ROLE

How did AI help?

OpenAI says the model ran with limited human supervision OpenAI post ↗. People still played a part: researchers sometimes suggested retrying strategies that had worked before, asked for proofs to be expanded after expert feedback, used ChatGPT to help with checking and formatting, and for some problems chose the best of several attempts OpenAI post ↗. The mathematical arguments themselves are presented as the model's work.

10problems attempted
5attempts OpenAI rates likely correct
1attempt OpenAI now believes is wrong

Figures are OpenAI's own assessment from its February 20, 2026 post OpenAI post ↗.

OpenAI was open about the limits. It called the effort a fast sprint and said its process was not as clean as a properly controlled evaluation OpenAI post ↗. That matters because First Proof organiser Lauren Williams asked how to judge how much is human and how much is AI once humans are involved Scientific American ↗. A Google team running a different system reported that its agent solved 6 of the 10 problems on its own by majority expert assessment, which shows that the scores of different labs used different rules arXiv preprint ↗.

03 · THE POSSIBILITIES

Which fields could this affect?

The immediate value is a public record of how an AI model performs on fresh research problems; wider uses are possible later, and these connections are our assessment.

Relevant now

Testing AI in research mathematics

First Proof gives a model for fair testing: unpublished problems, known answers and expert graders. The organisers have since run further batches with formal grading First Proof site ↗.

Explore science
Relevant now

Mathematical research practice

Mathematicians can read the attempts and the official solutions side by side OpenAI post ↗ Organiser solutions ↗. That helps them judge where AI proofs are strong and where they hide gaps.

Explore science
Possible future use

AI research assistants

If models become reliable on lemmas like these, they could help researchers with smaller steps inside bigger projects. This report does not show that reliability yet.

Explore software
04 · THE EVIDENCE

What has been checked?

The evidence is a set of published proof attempts, the company's own assessment, organiser solutions and news reporting, not a peer reviewed evaluation. Leapscope reviewed these sources; we did not repeat the experiments.

Shown so far

  • OpenAI published attempts for all ten First Proof problems and described the human feedback and selection involved OpenAI post ↗.
  • The organisers released their own solutions and commentary so experts can compare answers Organiser solutions ↗.
  • OpenAI withdrew its claim on problem 2 after official commentary and community analysis OpenAI post ↗.

Still unknown

  • Whether every attempt OpenAI rates as likely correct survives full expert checking; Nature reported the trial results were not officially verified Nature news ↗.
  • How much of each final proof depended on human choices, such as picking the best of several attempts OpenAI post ↗.
  • How the unreleased model would perform under a strictly controlled, no-help test OpenAI post ↗.

Evidence status: Proof attempts. Stage: Claim. Proof attempts shared publicly; independent checking is not complete.

05 · WHAT COMES NEXT

From sprint to fair test

  1. Finish the expert checks.Specialists in each field need to confirm or reject the remaining attempts.
  2. Run a controlled test.OpenAI says it hopes to discuss a more rigorous evaluation framework with the organisers OpenAI post ↗.
  3. Compare under the same rules.Results from different labs only mean much when the human help allowed is the same.

This is our suggested way to follow the story, not a promised timetable.

Can I use it today?

You can read OpenAI's post and its preprint of all ten attempts, and the organisers' solutions and comments OpenAI post ↗ Organiser solutions ↗. The model that wrote the attempts is internal and not available to the public OpenAI post ↗.

06 · QUICK QUESTIONS

A few things you might be wondering

Did OpenAI's AI solve ten research problems?

No. It produced attempts for all ten. OpenAI itself rates about five as likely correct and now believes one of its earlier favourites, problem 2, is wrong OpenAI post ↗.

Did the AI work completely on its own?

Not entirely. People suggested retries, asked for clearer proofs and sometimes picked the best of several attempts OpenAI post ↗. OpenAI said the process was not a properly controlled evaluation OpenAI post ↗.

Is this the same as being published in a journal?

No. These are proof attempts plus the company's own assessment. Nature later reported that the February trial results were not officially verified by the First Proof team Nature news ↗.

THE READING LIST

Go straight to the sources

Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.

01
First Proof submissionsOpenAI · February 20, 2026

OpenAI's explanation of its ten proof attempts, the human role and its own correctness estimates.

02
AI just got its toughest math test yet. The results are mixedScientific American · February 2026

News report on the challenge, how it was judged and comments from the organisers and other mathematicians.

03
First ProofFirst Proof initiative · website

The organisers' site with the board, timeline of batches and links to solutions.

04
Aletheia tackles FirstProof autonomouslyarXiv preprint · February 24, 2026

A Google team's report on a different AI agent's results on the same ten problems.

05
Humans outperform AI at this highly rigorous mathematics testNature news · June 12, 2026

Report on the second First Proof batch, noting the February trial was not officially verified.

06
First Proof Solutions and CommentsCowles Foundation discussion paper · April 2026

The organisers' own solutions and their discussion of AI responses.

ONE DISCOVERY LEADS TO ANOTHER

Keep following the possibilities.

AI × MATHEMATICS

722 manuscripts. A new scale of AI mathematics.

AI × MATHEMATICS

Gold medal standard in mathematics

FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.