What happened?
In February 2025, Google introduced the AI co-scientist, a system of several AI agents built on Gemini 2.0 that suggests new hypotheses and research plans from a goal written in plain language Google blog ↗. A hypothesis is a proposed explanation that can be tested by experiment. Google and partner labs tested some of its ideas in three areas of biomedicine arXiv paper ↗.
The first test was drug repurposing, which means finding new uses for existing medicines. The system proposed drugs for acute myeloid leukemia (AML), a blood cancer, and lab experiments showed some of them reduced the survival of cancer cells in dishes at doses relevant to patients Google blog ↗. In a second test, it suggested targets for liver fibrosis, a scarring disease, and Stanford collaborators reported anti fibrotic activity in lab grown mini livers called organoids Google blog ↗.
The third test was a check against known answers. Scientists at Imperial College London asked how a group of bacterial gene carriers called cf-PICIs spread across different bacterial species, a question they had solved in the lab but not yet published Google blog ↗. The system's top ranked idea, that cf-PICIs borrow tails from many different viruses that infect bacteria, matched their finding Cell paper ↗. That work was later published in the journal Cell Cell paper ↗.
How do the agents divide the work?
Supervisor
Breaks the research goal into a plan and hands tasks to the other agents Google blog ↗.
Generate and reflect
One agent writes hypotheses and another reviews them, like a peer reviewer Google blog ↗.
Rank and evolve
Ideas compete in tournaments, and the strongest are combined and refined over many rounds Google blog ↗.
In a small number of cases, ideas from the system held up in real lab experiments. That points to AI helping scientists choose which experiments to run first.
Leapscope interpretation of the reported result.How did AI help?
Scientists set the research goal and could add their own ideas or feedback in plain language Google blog ↗. The AI agents searched, debated and ranked hypotheses, using extra computing time to improve them over many rounds arXiv paper ↗. Human researchers then chose which ideas to test and did all of the laboratory work Google blog ↗. In the cf-PICI case, the scientists noted that some of the AI's other top hypotheses also opened new research directions in their labs Cell paper ↗.
Figures from Google's announcement Google blog ↗ and the research paper arXiv paper ↗.
The evidence comes mainly from Google and its partner labs, and the expert ratings used a small sample, which the authors acknowledge Google blog ↗. Outside researchers interviewed by TechCrunch were doubtful: one pathologist said the AML results were too vague to judge, and others said the hard part of science is designing and running experiments, which the system does not do TechCrunch ↗. Google lists better literature review, fact checking and larger expert evaluations as work still needed Google blog ↗.
Which fields could this affect?
The immediate value is help choosing experiments; uses in medicine are further off. These connections are our assessment.
Biomedical research
Selected research groups can apply for access through a Trusted Tester program Google blog ↗. The tool suggests ideas; labs still have to test them.
Explore scienceMicrobiology
The cf-PICI test showed the system could reach a known but unpublished answer Cell paper ↗. It is one example, not proof that it works on every question.
Explore scienceDrug repurposing
The AML candidates worked against cancer cells in dishes Google blog ↗. Cell experiments are far from showing a treatment works in people.
Explore healthcareNew treatments
The liver fibrosis results came from lab grown organoids Google blog ↗. No treatment for patients has been demonstrated.
What has been checked?
The evidence is a research preprint, a developer announcement and a peer reviewed Cell paper on the bacterial case, plus outside expert reaction. Leapscope reviewed these sources; we did not repeat the experiments.
Shown so far
- Suggested AML drugs reduced cancer cell survival in lab dishes Google blog ↗.
- Suggested liver fibrosis targets showed activity in human liver organoids Google blog ↗.
- The top hypothesis on cf-PICIs matched an unpublished lab finding, later published in Cell Cell paper ↗.
Still unknown
- How well it performs when tested independently, outside Google and its partners.
- How often its ideas fail, since mainly successes were reported.
- Whether any suggestion leads to a treatment that works in people.
Evidence status: Research demonstration. Stage: Claim. Results come from the developer and partner labs. Wider testing is still needed.
From hypothesis to proof
This is our suggested way to follow the story, not a promised timetable.
Can I use it today?
Not for most people. Google offers access to research organisations that apply to its Trusted Tester program Google blog ↗. Anyone can read the paper describing how the system works arXiv paper ↗.
A few things you might be wondering
Did the AI solve a 10 year mystery in days?
It proposed, in days, the same explanation an Imperial College team had found through years of lab work but not yet published PsyPost ↗. It matched a known answer; the humans did the discovery in the lab Cell paper ↗.
Does it replace scientists?
No. Scientists set the goals, choose which ideas to test and run the experiments Google blog ↗. Some outside researchers say generating ideas is not the main bottleneck in science TechCrunch ↗.
Has it produced a new medicine?
No. The drug results are from cells and organoids in the lab, not patients Google blog ↗.
Go straight to the sources
Checked Oct 8, 2026. The first source is the original announcement or research. Later sources add independent context; background pages do not validate the result on their own.
01Google's description of the system, its agents and three lab validated examples.
The technical paper by Gottweis and colleagues on the method and evaluations.
Outside researchers explain their doubts about the tool and Google's evidence.
Reports the published Cell papers on the cf-PICI case, with caveats.
Penadés, Gottweis and colleagues compare the AI's hypotheses with their own lab findings.