Stanford researchers have built an experimental AI-powered research organization that looks less like a chatbot and more like a virtual biotechnology company.
Called Virtual Biotech, the system used more than 37,000 AI clinical-trialist agents to analyze outcomes from 55,984 clinical trials. The agents helped researchers identify biological features associated with better drug-development outcomes and investigated therapeutic questions involving lung cancer and ulcerative colitis.
The headline number is impressive. But the most interesting part of the research is not simply that Stanford ran 37,000 AI agents.
It is what those agents were organized to do.
One part of the project has also created a potentially misleading headline. Stanford’s AI independently built a scientific case for targeting B7-H3 with an antibody-drug conjugate (ADC) in lung cancer. Separately, Daiichi Sankyo had already developed ifinatamab deruxtecan, a B7-H3-directed ADC that is now being jointly developed with Merck.
Later clinical results supported that same therapeutic direction.
But Merck did not clinically validate a drug molecule invented by Stanford’s AI agents.
That distinction matters.
The bigger story is whether coordinated teams of AI agents can take over larger parts of the reasoning, evidence gathering, analysis and coordination that normally happen across an entire biotech research organization.
For a broader view of where AI systems are heading across industries, see our Top 100 Best AI Tools in 2026.
Quick Answer: What Did Stanford’s 37,000 AI Agents Do?
Stanford’s Virtual Biotech is a multi-agent AI framework designed to resemble a therapeutic research organization.
For one large-scale project, it deployed more than 37,000 task-specific clinical-trialist agents to process evidence from 55,984 clinical trials. The resulting analysis found associations between cell-type-specific drug targets and better clinical-development outcomes.
Here are the headline figures:
| Finding | Reported Result |
|---|---|
| Clinical trials analyzed | 55,984 |
| Clinical-trialist agents deployed | 37,000+ |
| Phase I to Phase II progression | 40% more likely |
| Reaching Phase IV or market | 48% more likely |
| Adverse-event rates | 32% lower |
| Biomedical tools available | 100+ |
The biological signal involved cell-type specificity.
According to the research, drugs targeting genes with more cell-type-specific expression were 40% more likely to progress from Phase I to Phase II, 48% more likely to reach Phase IV or market, and associated with 32% lower adverse-event rates.
These findings are important, but they need careful interpretation.
They are retrospective associations identified by analyzing existing clinical trials. They do not prove that selecting a cell-type-specific target will automatically make a future medicine more likely to succeed.
Instead, the result identifies a biological feature that may deserve more attention when researchers prioritize potential drug targets.
What Is Stanford’s Virtual Biotech?
Virtual Biotech is designed around a simple idea:
Complex scientific research may work better when AI is organized into specialized teams instead of asking one general-purpose model to solve everything.
At the top of the system is a virtual Chief Scientific Officer, or CSO.
The CSO receives a scientific question, breaks it into smaller research tasks and delegates those tasks to specialized scientist agents.
The research describes eight scientist agents spread across four research divisions, together with additional roles in the CSO’s office. Task-specific agents can then be created at much larger scale when the workload requires it.
This structure matters because drug development is not a single problem.
Researchers may need to consider:
- human genetics
- disease biology
- functional genomics
- single-cell data
- drug-target interactions
- chemistry
- safety evidence
- clinical trial outcomes
- biomarkers
- therapeutic modality
A single AI model can potentially examine all of these areas, but keeping every discipline, dataset and scientific tradeoff inside one reasoning process becomes difficult.
Virtual Biotech instead attempts to reproduce something closer to a cross-functional scientific organization.

Why Did Stanford Need 37,000 AI Agents?
The 37,000 figure can create the wrong mental picture.
Stanford did not create 37,000 permanent digital scientists and put them into one enormous virtual research meeting.
The core Virtual Biotech organization is much smaller.
The 37,000-plus clinical-trialist agents were spawned for a large parallel data-processing task.
Stanford Medicine describes these agents as being used to read and analyze clinical-trial results from sources including research papers, trial registries and press releases.
In other words, the number reflects massive parallel task execution.
A smaller leadership structure decides what needs to be investigated. Thousands of task-specific agents can then work on individual pieces of evidence.
That architecture is relevant far beyond biotechnology. Multi-agent systems are increasingly being explored for workflows in which one AI coordinates specialized agents, tools and business processes. Aitoza’s guide to the Top 25 AI Tools for Business Automation in 2026 looks at the same broader shift toward AI-driven workflow orchestration.
How Virtual Biotech Gives AI Agents Scientific Tools
An AI model cannot conduct serious biomedical research from language ability alone.
It needs access to scientific tools and structured evidence.
According to the Virtual Biotech paper, its agents can use more than 100 customized tools and analytical resources covering areas such as statistical genetics, functional genomics, chemistry, disease biology and clinical evidence.
The research environment includes information covering tens of thousands of biological targets and diseases, millions of genetic credible sets, more than 100 million single-cell profiles, thousands of drugs, adverse-event records and large-scale perturbation datasets.
The key advantage is not simply having access to a huge amount of information.
It is being able to examine the same therapeutic question through multiple scientific perspectives.
A genetics agent may find limited inherited genetic support for a target.
A single-cell agent may identify strong tumor-specific expression.
Another agent may investigate safety signals.
Another may examine whether the target is better suited to an antibody, small molecule or other therapeutic approach.
The CSO agent can then integrate those findings into a broader recommendation.
That is fundamentally different from asking a chatbot one question and accepting its first answer.
What Did the AI Find Across 55,984 Clinical Trials?
One of the largest Virtual Biotech demonstrations asked:
What properties of drug targets are associated with better clinical outcomes?
Answering this required connecting trial outcomes with genomic and biological characteristics of the targets those medicines were designed to affect.
More than 37,000 agents helped structure information from 55,984 clinical trials.
The analysis highlighted cell-type specificity.
40% Higher Phase I to Phase II Progression
The study reported that drugs targeting cell-type-specific genes were 40% more likely to progress from Phase I to Phase II than the comparison group.
That does not mean cell-type specificity guarantees success.
Clinical development can fail because of efficacy, toxicity, dosing, patient selection, manufacturing, commercial strategy and many other factors.
Instead, the finding suggests that cell-type specificity may be a useful signal when researchers are choosing which targets deserve further investment.
48% Higher Likelihood of Reaching Phase IV or Market
The analysis also reported a 48% increase in the likelihood of reaching Phase IV or market for drugs aimed at cell-type-specific genes. Stanford later highlighted this result during its 2026 Health AI Week coverage.
Again, this is an association, not proof of causation.
A prospective experiment would be considerably stronger: let the AI use this information to select new targets before outcomes are known and then follow those programs through laboratory and clinical development.
32% Lower Adverse-Event Rates
The same analysis found 32% lower adverse-event rates among drugs targeting cell-type-specific genes.
One possible explanation is that a target concentrated in particular cell types could allow a therapy to act more selectively.
But the current result should be treated as hypothesis-generating evidence rather than a universal drug-design rule.
The B7-H3 Lung Cancer Case Explained
The second major Virtual Biotech example focused on B7-H3, also known as CD276, as a therapeutic target in lung cancer.
Rather than relying on one evidence source, the AI system examined B7-H3 through several scientific lenses, including genetics, single-cell data, spatial data, clinicogenomic information, existing clinical evidence and treatment modality.
Interestingly, the genetic evidence alone was not especially strong.
The AI did not automatically reject the target.
Instead, it reasoned that cancer targets can sometimes be supported by tumor-specific biology even when inherited genetic evidence is limited.
After integrating the wider evidence, Virtual Biotech supported B7-H3 as a therapeutic target and proposed an antibody-drug conjugate strategy.
Why an Antibody-Drug Conjugate?
An antibody-drug conjugate combines an antibody that recognizes a target on cells with a potent drug payload.
The goal is to use the antibody to direct the payload toward cells expressing the target.
For Virtual Biotech, B7-H3’s cell-surface expression and tumor biology made an ADC a logical therapeutic modality.
The important point is that the recommendation emerged from the integration of several forms of evidence rather than a single text-generation step.
Did Merck Validate a Stanford AI-Designed Cancer Drug?
No, not in the literal sense.
This is the most important clarification in the article.
Stanford’s Virtual Biotech did not invent ifinatamab deruxtecan.
Ifinatamab deruxtecan, also known as I-DXd, was discovered by Daiichi Sankyo and is being jointly developed with Merck. It is a B7-H3-directed DXd antibody-drug conjugate.
Stanford’s AI independently reached the conclusion that B7-H3 plus an ADC approach could be promising in lung cancer.
That therapeutic direction was later supported by clinical evidence involving an independently developed drug targeting the same biology.
So the accurate takeaway is:
Stanford’s AI independently identified and supported a therapeutic strategy that later clinical results also supported. It did not design the exact Merck and Daiichi Sankyo drug.
That is still scientifically interesting.
It means a multi-agent AI system was able to assemble different types of biomedical evidence and reach a therapeutic conclusion that was consistent with later clinical evidence.
What Did the IDeate-Lung01 Trial Find?
The relevant drug is ifinatamab deruxtecan, a B7-H3-directed ADC being developed by Daiichi Sankyo and Merck.
In the Phase II IDeate-Lung01 study, 137 patients receiving the 12 mg/kg dose were included in the combined analysis.
The confirmed objective response rate was 48.2%.
The reported disease-control rate was 87.6%, while median progression-free survival was 4.9 months and median overall survival was 10.3 months.
These clinical results do not prove that Stanford’s entire Virtual Biotech system can reliably discover successful drugs.
They provide external support for one therapeutic direction produced by the AI analysis.
What Is the FDA Status?
As of August 11, 2026, ifinatamab deruxtecan has not received the U.S. approval covered by its current small-cell lung cancer application.
The FDA accepted the Biologics License Application and granted it Priority Review in April 2026. The regulatory action date is October 10, 2026.
The FDA had previously granted the drug Breakthrough Therapy Designation for previously treated extensive-stage small cell lung cancer.
This status should be updated after October 10, 2026, or earlier if the FDA announces a decision before that date.
Virtual Biotech Also Investigated a Failed Drug Program
Virtual Biotech was not used only to identify promising therapeutic targets.
Stanford also asked the system to investigate a terminated Phase II ulcerative colitis program targeting OSMRβ.
The agents analyzed the failed program and proposed that a precision-medicine gap involving patient selection and biomarkers may have contributed to the outcome.
This is a potentially important application of agentic AI.
Drug companies learn enormous amounts from failed trials.
A program can fail because the underlying therapeutic hypothesis is wrong. But failure can also come from selecting the wrong patients, choosing weak biomarkers, using an unsuitable dose or testing a treatment at the wrong disease stage.
An AI system capable of systematically revisiting failed trials could help researchers determine whether a program should truly be abandoned or whether the original development strategy deserves another look.
Virtual Biotech Grew Out of Stanford’s Earlier Virtual Lab

Virtual Biotech builds on Stanford’s earlier Virtual Lab research.
The Virtual Lab used an AI principal investigator to coordinate specialized AI scientist agents while human researchers provided high-level guidance.
Stanford applied that system to the design of nanobodies targeting newer variants of SARS-CoV-2.
The AI team developed a computational workflow using tools including ESM, AlphaFold-Multimer and Rosetta and designed 92 new nanobody candidates.
This time, researchers went beyond computational analysis.
The candidates were physically tested.
Several were functional, and two showed improved binding to newer JN.1 or KP.3 variants while maintaining strong binding to the ancestral viral spike protein. The work was peer-reviewed and published in Nature in 2025.
That distinction is important.
The earlier Virtual Lab project demonstrates that AI-agent-generated scientific designs can move from software into real laboratory experiments.
It does not prove that Virtual Biotech can autonomously develop a complete medicine.
But it shows that Stanford’s agentic-science program is not limited to generating plausible scientific text.
Can AI Agents Make Drug Discovery Cheaper?
One striking figure in the Virtual Biotech research concerns computational cost.
The paper reports that the multi-modal B7-H3 analysis used only a relatively small amount of model API resources compared with the cost of conventional pharmaceutical development.
That does not mean an AI company can develop a cancer medicine for the price of a few model calls.
A real drug-development program still requires expensive physical and human infrastructure, including laboratory experiments, manufacturing, toxicology studies, regulatory work and clinical trials.
The more realistic opportunity is earlier in the decision-making process.
If AI agents can eliminate weak targets sooner, uncover overlooked evidence, identify better biomarkers or suggest stronger patient-selection strategies before expensive experiments begin, the savings could become substantial.
The value would come from making better research decisions earlier, not from eliminating the physical cost of biology.
What Stanford’s Virtual Biotech Still Cannot Do
The scale of the project makes it tempting to jump from “37,000 AI agents analyzed clinical trials” to “AI can replace pharmaceutical research teams.”
The evidence does not support that conclusion.
1. The Virtual Biotech Study Is Still a Preprint
The Virtual Biotech research was posted to bioRxiv in February 2026 and is currently a preprint rather than a completed peer-reviewed journal publication.
That does not make the findings unimportant.
It does mean they should be interpreted as ongoing research rather than settled scientific evidence.
Stanford’s earlier Virtual Lab nanobody work is different because that research completed peer review and appeared in Nature.
2. Much of the Evidence Is Retrospective
The 55,984-trial analysis looks backward at trials whose outcomes are already known.
Retrospective analysis is useful for finding patterns.
Prospective prediction is harder.
The stronger test would be for Virtual Biotech to select a new drug target, modality or biomarker strategy before the result is known, then test that prediction experimentally.
3. AI Agents Can Share the Same Blind Spots
Adding more agents does not automatically create more independent intelligence.
If agents depend on similar foundation models, datasets or assumptions, they may repeat the same error at scale.
Multiple AI agents agreeing with one another is therefore not equivalent to several independent human research teams confirming the same result.
4. Wet-Lab Validation Still Matters
AI agents can read scientific literature, query databases, execute code, compare evidence and suggest experiments.
Biology still has to be tested physically.
A target that looks excellent computationally can fail in living systems.
A drug can hit the right target and still cause unacceptable toxicity.
A biomarker that looks convincing in retrospective data can fail in a larger patient population.
AI does not remove those realities.
Why 37,000 Agents May Matter Less Than the Organizational Design
The most important idea behind Virtual Biotech may not be its headline number.
For years, much of AI progress focused on building a more capable model and asking that model increasingly difficult questions.
Stanford’s experiment points toward another possibility:
Instead of only making one AI smarter, organize multiple AIs more effectively.
Drug research is naturally suited to this approach because it is highly interdisciplinary.
A geneticist may understand whether human genetics supports a target.
A cancer biologist may understand what that target does inside a tumor.
A chemist may know whether it is druggable.
A safety scientist may identify potential liabilities.
A clinician may recognize that the correct patient population is much narrower than initially assumed.
Human biotechnology companies solve this problem through departments, meetings, reviews, project teams and management structures.
Virtual Biotech attempts to reproduce parts of that coordination in software.
Instead of one chatbot generating a final response, a leadership agent can delegate tasks, specialist agents can investigate different evidence sources, reviewers can challenge conclusions and the organization can integrate the findings.
For examples of how agentic workflows are also being applied to enterprise processes, Aitoza’s UiPath AI overview covers AI-driven automation outside biomedical research.
The Bigger Shift Toward AI Agents in Science
Virtual Biotech is part of a broader transition from AI systems that simply answer questions toward systems capable of pursuing longer, multi-step goals.
Stanford researchers and other scientific teams are increasingly studying agents that can select tools, analyze data, coordinate tasks, generate hypotheses and revise their work based on new information. Stanford’s 2026 AI Index also identifies multi-agent AI and AI-enabled biological research as important areas of development in medicine.
That creates a possible progression:
Chatbot → research assistant → specialized agent → multi-agent research team → AI-native scientific organization
Virtual Biotech is an experiment in that final category.
The difficult question is no longer simply whether AI can produce a scientifically plausible answer.
It is whether an AI organization can repeatedly produce useful, traceable and experimentally valid scientific decisions.