top of page

How Is AI in Biology Transforming Research, Healthcare, and Drug Discovery? (2026)

4 minutes ago
25 min read
AI in biology with DNA, medical imaging, lab research, and drug discovery.

In October 2024, the Nobel Prize in Chemistry honored two linked achievements: predicting a protein's three-dimensional shape from its amino acid sequence, and designing entirely new proteins by computer [3]. Since then, AI in biology has spread into DNA regulation, single-cell analysis, pathology and drug design. On 2026-09-09, a developer reported dosing the first patient in a Phase III trial of a drug whose target and molecule both came from generative AI [14]. That progress is real, and it is easy to overstate. This guide separates what AI in biology has demonstrated from what it has only promised, so researchers, clinicians, founders and buyers can use it with clear eyes.


TL;DR


  • What it is. AI in biology uses machine learning to find patterns in biological data and to predict or design proteins, cells, molecules and experiments. Its best role today is narrowing which experiments to run, not replacing them.

  • Strongest evidence. Protein structure prediction is mature and shared in the 2024 Nobel Prize in Chemistry [3]. Protein design is a separate task, with experimentally tested binders reported for RFdiffusion [4].

  • Drug discovery. AI-designed molecules have reached human trials. A Phase 2a trial of rentosertib was published in 2025 [13], and its developer reports a Phase III start on 2026-09-09 [14]. The drug is still investigational, not an approved medicine.

  • Benchmark caution. In an independent test, five single-cell foundation models and two other deep-learning models did not beat simple baselines at predicting perturbation effects [9].

  • Buying advice. Define one narrow problem, compare against a simple baseline, and require validation that matches the claim.

  • Key caveat. A faster computer screen does not mean a faster approval. Efficacy and safety are still decided by experiments and clinical trials [16][17].


How is AI transforming biology, healthcare, and drug discovery?


AI is helping biologists predict protein structures, interpret DNA variants, analyze cells and images, and shortlist drug candidates, which narrows expensive experiments. The evidence is strongest for structure prediction and weakest for patient benefit. Faster computation does not automatically mean faster approval, because experiments and clinical trials still decide whether a medicine works.

What is the biggest barrier to AI delivering real-world impact in biology?

  • 0%Poor or fragmented biological data

  • 0%Wet-lab validation

  • 0%Clinical evidence

  • 0%Model reliability and generalization


Table of Contents


  1. What Is AI in Biology?

  2. Why AI and Biology Are Converging Now

  3. How AI Learns From Biological Data

  4. How AI Is Transforming Genomics and Precision Biology

  5. How AI Is Changing Single-Cell, Spatial and Systems Biology

  6. How AI Is Transforming Protein Structure and Protein Design

  7. How AI Is Changing Microscopy, Pathology and Biological Imaging

  8. How AI Is Transforming Drug Discovery and Development

  9. What AI Means for Healthcare and Precision Medicine

  10. Real-World Examples of AI in Biology

  11. What Can AI in Biology Actually Do Today, and What Can't It Do?

  12. Limitations, Risks and Scientific Challenges

  13. Regulation, Governance and Responsible AI in Biology

  14. The Commercial Landscape: How Organizations Use AI Biology Platforms

  15. Build vs Buy vs Partner: How Should a Biotech or Pharma Team Decide?

  16. How to Evaluate an AI Biology Platform or Vendor

  17. How to Implement AI in a Biology or Drug-Discovery Organization

  18. Where AI in Biology Is Heading Next

  19. So, How Is AI in Biology Transforming Research, Healthcare, and Drug Discovery?

  20. Frequently Asked Questions

  21. Key Takeaways

  22. Actionable Next Steps

  23. Glossary

  24. Sources & References


What Is AI in Biology?


AI in biology is the use of machine learning to find patterns in biological data and to predict, generate or rank biological objects, such as protein structures, the effects of DNA variants, cell states, drug-like molecules and clinical signals. It is a toolkit applied across the life sciences, not one technology.


AI, machine learning, deep learning, generative AI and foundation models


These terms nest inside one another. Artificial intelligence is the broad goal of building systems that perform tasks that seem to need intelligence. Machine learning is the part of AI that learns rules from examples instead of following hand-written instructions. Deep learning uses many-layered neural networks and drives most modern biology models. Generative AI creates new data, such as sequences, structures or images. Foundation models are large models pretrained on broad data and then adapted to many tasks.


Biology now has an example of each idea. Geneformer was pretrained on about 30 million single-cell transcriptomes [8]. UNI was pretrained on more than 100 million image patches from over 100,000 pathology slides [12]. Evo 2 was trained on 9 trillion DNA base pairs from all domains of life [7]. AlphaFold 3 and RFdiffusion use diffusion methods, which learn to remove noise step by step until a structure emerges [2][4].


Computational biology versus AI


Computational biology is the wider field that uses statistics, simulation and computing to study living systems. AI is one set of tools inside it. Older methods still beat newer AI on some tasks, which is why baselines matter later in this guide [9][10].


What AI in biology is not


It is not an autonomous scientist. In the Virtual Lab study published in Nature, language-model agents proposed a nanobody design pipeline, but human researchers gave feedback and tested 92 designed nanobodies in the laboratory [18]. The pattern is typical: people design, curate and validate, and models speed up the search.


Why AI and Biology Are Converging Now


Three ingredients arrived together: large biological datasets, model architectures that scale, and growing capacity to test predictions in the laboratory.


Data came first. Before AlphaFold, experiments had solved the structures of only around 100,000 unique proteins, a small fraction of the billions of known sequences [1]. Large cell atlases and public genomics programs now supply the millions of cells and thousands of measurements that models such as Geneformer and AlphaGenome learn from [6][8].


Architecture and computing power came next. Transformers and diffusion models, first proven on text and images, transferred to sequences and molecules. AlphaFold 2 was validated in the blind CASP14 assessment, where its accuracy was competitive with experiments in a majority of cases [1]. The third ingredient is closing the loop. Automated experiments let teams test many model outputs quickly, so predictions get checked instead of merely admired.


How AI Learns From Biological Data


AI models turn biological measurements into numbers, learn patterns from examples, and then predict labels or generate new examples. The quality and labeling of the data usually limit the result more than the model does.


Sequences are read as strings of letters, much like text. Structures are stored as 3D coordinates or atom graphs. Images are grids of pixels. Single-cell data are large tables of gene activity per cell. Models learn an embedding, a compact numerical summary of each item, and use it for tasks such as clustering cells or scoring variants. Labels are measured outcomes, such as whether a variant causes disease or a compound binds a target. Labels are expensive, so many models use self-supervised pretraining on unlabeled data, followed by fine-tuning on small labeled sets.


Table 1 summarizes the main data types, tasks and maturity levels.


Biological data

AI task

Example

Potential value

Maturity and limits

DNA and RNA sequence

Predict regulatory activity and variant effects

AlphaGenome [6]; Evo 2 [7]

Prioritize non-coding variants

Research use; a prediction is not a diagnosis

Protein sequence and structure

Predict structures and complexes; design binders

AlphaFold 3 [2]; RFdiffusion [4]

Faster target and binder work

Prediction is mature; designs need lab tests

Single-cell and omics tables

Cell embeddings; perturbation prediction

Geneformer [8]

Map cell states

Mixed benchmarks [9][10]

Microscopy and slide images

UNI [12]

Review images at scale

Domain shift; clinical validation needed

Molecular graphs

Predict properties; generate molecules

Rentosertib program [13]

Design candidate drugs

Benefit unproven until trials succeed

Mixed data types

Joint models across scales

Virtual cell proposal [11]

Simulate cells and tissues

Early stage, largely a vision


Two rules of thumb follow. First, match the evidence to the claim: a strong benchmark score on public data does not prove that a model will work on your samples. Second, ask what the labels really measure, because a model trained on cell-line data may not transfer to patient tissue.


How AI Is Transforming Genomics and Precision Biology


AI helps genomics by predicting which DNA changes are likely to alter gene function, so researchers and clinical geneticists can prioritize variants for follow-up. These predictions are hypotheses about mechanism and risk, not diagnoses.


Variant interpretation: AlphaMissense


A missense variant swaps one amino acid in a protein, and most observed variants have unknown clinical meaning. AlphaMissense adapts AlphaFold and is fine-tuned on human and primate variant frequency data. In Science (2023-09-22), its authors reported state-of-the-art results across genetic and experimental benchmarks and classified 89% of missense variants as likely benign or likely pathogenic [5]. A classification like this can rank variants for review. It does not replace clinical evidence.


Gene regulation: AlphaGenome and DNA language models


Much of the genome does not code for proteins, yet it holds regulatory switches that variants can disturb. AlphaGenome reads 1 megabase of DNA and predicts thousands of functional tracks at up to single-base resolution. Published in Nature on 2026-01-28, it matched or exceeded the strongest external models in 25 of 26 variant-effect evaluations, and it recapitulated mechanisms of clinically relevant variants near the TAL1 oncogene [6]. Evo 2, a DNA language model trained on 9 trillion base pairs, extends this style of modeling toward design as well as prediction [7].


The limits are practical. These models learn correlations from measured data, much of it from a limited set of cell types and populations. A predicted change in a chromatin track is not a proven change in disease risk, and association is not causation. Laboratory tests and expert clinical curation stay essential.


How AI Is Changing Single-Cell, Spatial and Systems Biology


Single-cell and spatial methods measure thousands of genes in individual cells or tissue positions. AI helps annotate cell types, integrate datasets and predict how cells respond to changes. The current evidence is mixed: large models help in some tasks, but simple baselines often match them.


Geneformer, published in Nature in 2023, was pretrained on about 30 million single-cell transcriptomes. After fine-tuning, it improved predictions on tasks related to chromatin and gene networks, and it nominated candidate therapeutic targets for cardiomyopathy [8]. Candidate targets are leads for testing, not validated treatments.


What independent benchmarks found


Ahlmann-Eltze, Huber and Anders (Nature Methods, 2025-08-04) compared five foundation models and two other deep-learning models against deliberately simple baselines for predicting transcriptome changes after single or double gene perturbations. None outperformed the baselines [9]. Kedzierska and colleagues (Genome Biology, 2025) found that Geneformer and scGPT, used without further training, could be outperformed by simpler methods in some cases [10]. These results cover specific tasks and datasets. They do not prove such models are useless. They show that any claimed advantage must be measured against a strong simple baseline, such as an additive model or highly variable genes.


Virtual cells and systems biology


The virtual cell is a research vision, not a finished product. In Cell (2024-12-12), Bunne and colleagues described an AI virtual cell as a multi-scale, multimodal model that could simulate molecules, cells and tissues, and they stressed the need for shared data, evaluation strategies and community standards [11]. Causal questions, such as what happens when a gene is switched off, require perturbation experiments, because observational data mainly show correlations. Spatial biology adds tissue location, which raises both the value and the integration burden of the data.


How AI Is Transforming Protein Structure and Protein Design


AI protein structure prediction estimates the 3D shape of a protein or complex from its sequence. AI protein design works in the opposite direction, proposing new sequences or shapes for a desired function. These are different tasks with different levels of evidence.


Structure prediction: AlphaFold 2 and AlphaFold 3


AlphaFold 2 was validated in the blind CASP14 assessment and provided the first method that could regularly predict structures with atomic accuracy when no similar structure was known [1]. AlphaFold 3, published in Nature in 2024, predicts joint structures of complexes containing proteins, nucleic acids, small molecules, ions and modified residues. Its authors reported higher accuracy than docking tools for protein–ligand interactions, than nucleic-acid-specific predictors for protein–nucleic acid interactions, and than AlphaFold-Multimer v2.3 for antibody–antigen complexes [2]. The 2024 Nobel Prize in Chemistry recognized Demis Hassabis and John Jumper for protein structure prediction, and David Baker for computational protein design [3].


Protein design: RFdiffusion


RFdiffusion fine-tunes the RoseTTAFold network to denoise protein structures, which lets it generate new backbones for binders, symmetric assemblies and metal-binding proteins. The authors experimentally characterized hundreds of designs, and a cryo-electron microscopy structure of a designed binder bound to influenza haemagglutinin was nearly identical to the design model [4]. That is stronger evidence than a computed prediction, because an experiment confirmed the structure.


Two cautions apply. A confident predicted structure does not prove that a protein binds a target tightly, folds in cells or works as a drug. And design still needs expression, binding and stability assays, so wet-lab throughput often sets the pace.


How AI Is Changing Microscopy, Pathology and Biological Imaging


AI reads microscopy and pathology images to segment cells, classify tissue and extract biomarkers at a scale people cannot match. Research use is well established. Clinical use is narrower and depends on validation and regulatory clearance.


In the laboratory, deep learning segments cells, measures shape and marker changes, and supports phenotypic screens that compare how compounds change the appearance of cells. In pathology, UNI (Nature Medicine, 2024-03-19) was pretrained on more than 100 million images from over 100,000 stained tissue slides across 20 tissue types and was evaluated on 34 tasks, including classification of up to 108 cancer types [12].


Translation is the hard part. Staining methods, scanners and hospital workflows change how images look, so performance can fall at a new site. A 2026 medRxiv preprint that analyzed the FDA's list of AI-enabled devices counted 1,430 authorizations from September 1995 to December 2025, with 76.5% in radiology and only nine in pathology [26]. It is not yet peer reviewed. For clinical imaging in more depth, see our guide to AI in medical imaging.


How AI Is Transforming Drug Discovery and Development


AI now touches almost every stage of drug discovery, from choosing targets to designing molecules and planning trials. The evidence is strongest for early computational tasks and for producing molecules that pass early safety testing. It is weakest for showing that AI-discovered drugs work better in patients.


Table 2 maps the pipeline and separates what a tool can do from what the evidence supports.


Stage

AI application

Expected benefit

Evidence and maturity

Principal limitation

Target identification

Mine omics, literature and networks for disease drivers

Faster hypothesis generation

Emerging; few targets validated in patients

Correlation is not causation

Target validation

Predict perturbation effects; prioritize experiments

Fewer dead-end targets

Mixed; simple baselines often competitive [9]

Needs wet-lab confirmation

Hit discovery and virtual screening

Score compounds against structures

Cheaper first-pass filtering

Widely used; AlphaFold 3 reports gains over docking [2]

Scoring errors; assay artifacts

De novo molecule generation

Propose molecules with target properties

New chemical ideas

Tested clinically in a few programs [13]

Synthesis and potency must be proven

Lead optimization and ADMET

Predict potency, metabolism and toxicity

Fewer optimization cycles

Useful but data-dependent

Weak transfer to new chemistry

Biomarkers and stratification

Find patient subgroups from omics and images

Better trial enrichment

Emerging; needs prospective testing

Overfitting on small cohorts

Clinical trial design

Model sites, endpoints and recruitment

Fewer delays

Early and mostly retrospective

Acceptance depends on context of use [20]

Manufacturing and monitoring

Process control and safety-signal detection

More consistent quality

Emerging in regulated settings

Lifecycle validation needed [22]


What the clinical evidence shows


Jayatunga and colleagues (Drug Discovery Today, 2024) analyzed the clinical pipelines of AI-native biotech companies. They reported a Phase I success rate of 80–90%, higher than historic averages, and about 40% in Phase II on a limited sample, comparable to historic averages [16]. Read this carefully. Phase I mainly tests safety and drug-like behavior, so the result supports AI's ability to design molecules that behave like drugs, not that those drugs cure disease more often. The Phase II sample was small.


Rentosertib is the most advanced example. Its randomized, double-blind, placebo-controlled Phase 2a trial (Nature Medicine, 2025-06-03) enrolled 71 people with idiopathic pulmonary fibrosis for 12 weeks. The primary endpoint was safety: adverse events occurred in 72.2% to 83.3% of patients across the rentosertib arms and 70.6% on placebo [13]. Trade press reported a mean lung-function (FVC) change of +98.4 mL at 60 mg once daily versus −20.3 mL on placebo [15]. The trial was small and short, and it was not designed to prove efficacy. The developer reports that the first patient was dosed on 2026-09-09 in GENESIS-IPF-3, a 52-week placebo-controlled Phase III trial expected to enroll 320 people [14]. The drug remains investigational.


Faster is not the same as approved


Speed at the computer does not shorten the biology. Efficacy, safety, manufacturing and regulatory review remain, and one AI-linked candidate has already failed its efficacy endpoints (see BEN-2293 below).


Pros and cons at a glance


  • Pros: wider search of targets and chemistry, earlier filtering of weak candidates, better use of structural data, and more consistent decisions when models are validated.

  • Cons: an unproven effect on late-stage success, dependence on data quality, the cost of validation, and the risk of over-trusting model scores.


What AI Means for Healthcare and Precision Medicine


AI reaches patients through five routes that carry different levels of evidence. Most biology models in this guide sit in the first route.


  • Research utility: tools such as AlphaMissense and AlphaGenome help researchers rank variants for study [5][6].

  • Diagnostic assistance: image and pathology tools support readers. FDA-authorized AI devices cluster in radiology [26].

  • Treatment selection: biomarker models may guide therapy, but they need prospective validation.

  • Clinical decision support: software that recommends actions to clinicians needs its own validation and oversight.

  • Approved products: authorized devices exist, but an authorized device is not an approved AI-discovered medicine.


Precision medicine matches care to a person's genes, tumor or biomarkers. AI can help interpret genomic data and rare-disease variants, yet clinical use requires validated, regulated tools and clinician oversight. For a wider clinical view, see AI in healthcare: 15 real applications. This article is educational and is not medical advice. Consult a qualified clinician for personal health decisions.


Real-World Examples of AI in Biology


Each case below names the problem, the approach, what was demonstrated, why it matters and what stays uncertain.


AlphaFold: structure prediction at scale


Experiments had solved only around 100,000 unique protein structures against billions of known sequences [1]. AlphaFold used a neural network that builds in evolutionary and physical knowledge, reached atomic accuracy in CASP14 [1], and in version 3 modeled complexes with ligands and nucleic acids [2]. It earned a share of the 2024 Nobel Prize in Chemistry [3]. What remains uncertain is binding strength, function and design, which prediction alone does not settle.


AlphaMissense: ranking missense variants


Most missense variants have unknown clinical meaning. AlphaMissense fine-tunes an AlphaFold-derived model on population variant frequencies and classified 89% of missense variants as likely benign or likely pathogenic [5]. It helps prioritize variants. It is not a diagnostic test.


Rentosertib: from an AI-identified target to Phase III


The problem was idiopathic pulmonary fibrosis, a progressive lung disease. Developers used generative AI for both target and molecule, then ran a Phase 2a trial [13] and, by their report, started Phase III on 2026-09-09 [14]. It matters as the first end-to-end example at this stage. The open question is whether the early lung-function signal holds in a larger, longer trial.


BEN-2293: a counterexample


BenevolentAI reported on 2023-04-04 that its topical BEN-2293 was safe and well tolerated in a Phase IIa trial of 91 adults with mild-to-moderate atopic dermatitis, but did not achieve its secondary efficacy endpoints for itch and inflammation [17]. The lesson is that computational selection does not guarantee clinical efficacy.


The perturbation benchmark: an evidence check


Ahlmann-Eltze and colleagues tested whether single-cell foundation models could beat simple baselines at predicting perturbation effects, and none did [9]. It matters because it shows how independent benchmarking corrects hype. The results apply to the tasks tested and may change as models and data improve.


The Virtual Lab: AI agents plus human scientists


Swanson and colleagues (Nature, 2025) built a team of language-model agents, led by an agent principal investigator with human feedback, to design SARS-CoV-2 nanobodies. Of 92 designs tested in the laboratory, over 90% expressed as soluble proteins, and two showed distinctive binding to recent JN.1 and KP.3 spike variants [18]. It shows a workable human–AI loop. Two promising binders are early leads, not therapies.


What Can AI in Biology Actually Do Today, and What Can't It Do?


Today, AI in biology is established for structure prediction and research image analysis, useful but context-dependent for variant ranking and molecule design, and emerging or experimental for perturbation prediction, virtual cells and autonomous discovery. Table 3 groups capabilities by maturity without arbitrary scores.


Capability

Status

Evidence needed

Key caveat

Protein structure prediction

Established

Blind assessments and experimental structures [1]

Not the same as design or function

Image analysis in research

Established

Multi-site validation [12]

Performance shifts across sites

Variant effect ranking

Useful, context-dependent

Benchmarks plus clinical curation [5][6]

Predictions are not diagnoses

De novo protein design

Useful, context-dependent

Wet-lab binding and stability assays [4]

Experimental success rates vary

AI-designed small molecules

Emerging

Randomized clinical trials [13][17]

Efficacy unproven at scale

Perturbation prediction

Emerging

Beating simple baselines on new data [9]

Often matched by simple baselines

Virtual cells

Experimental

Prospective experimental validation [11]

Currently a research vision

Autonomous discovery agents

Experimental

Independent replication [18]

Human oversight is essential


Myths versus facts


  • Myth: Bigger biology models always win. Fact: Independent tests found simple baselines matched or beat several large models [9][10].

  • Myth: AlphaFold designs new drugs. Fact: It predicts structures, and design is a separate task [3][4].

  • Myth: AI-discovered drugs are proven to succeed more often. Fact: Phase I success is high, but efficacy evidence is limited [16][17].

  • Myth: AI will replace scientists. Fact: Current systems depend on human design, curation and validation [18].


Limitations, Risks and Scientific Challenges


The main limits are data quality, generalization, validation and governance, not raw model power. Each one can make a strong benchmark result misleading.


  • Data quality and bias: training data come from limited tissues, cell lines and populations, so models can fail for groups or conditions that were under-sampled.

  • Confounding and batch effects: technical differences between labs can look like biology, and models can learn the shortcut.

  • Generalization and distribution shift: performance often drops on data unlike the training data, such as a new hospital, assay or chemical series.

  • Benchmark leakage and reproducibility: if test data overlap training data, scores inflate. See our explainer on data leakage.

  • Interpretability and causation: a score without a mechanism is hard to audit, and observational data alone cannot show that a change causes an outcome.

  • Validation: computational, wet-lab, clinical and regulatory evidence are separate levels, and skipping one is a red flag.

  • Privacy, IP and data rights: patient-data permissions, licenses to training data and ownership of model outputs should be settled early.

  • Compute and cost: large models are costly to train and run, so ask whether a smaller model is enough.

  • Biosecurity and dual use: researchers showed that open protein-design software could produce variants of proteins of concern that evaded DNA-synthesis screening tools, and patches were then developed and deployed [19]. The response is governance and screening, not procedures.


Regulation, Governance and Responsible AI in Biology


Regulators expect AI used in drug development to be credible for a defined context of use, validated in proportion to risk and managed across its lifecycle. Most guidance is still draft or non-binding.


In January 2025, the FDA published draft guidance on AI used to produce information or data that supports regulatory decisions about a drug's safety, effectiveness or quality. It proposes a risk-based credibility assessment framework tied to a model's context of use [20][21]. As of 2026-09-30, the FDA guidance page still labels it draft guidance that is not for implementation [20].


On 2026-01-14, the FDA and EMA jointly released ten guiding principles of good AI practice in drug development: human-centric by design, a risk-based approach, adherence to standards, a clear context of use, multidisciplinary expertise, data governance and documentation, model design and development practices, risk-based performance assessment, life cycle management, and clear, essential information [22][23]. They are principles for developers to consider, not binding rules.


EMA's reflection paper on AI in the medicinal product lifecycle, adopted in September 2024, gives considerations for development, authorization and post-authorization use [24]. NIST's AI Risk Management Framework adds a voluntary structure built on Govern, Map, Measure and Manage functions [25]. See also our explainer on the NIST AI RMF.


Five ideas recur. Context of use is the exact role of a model in a decision. Risk-based validation demands more evidence for higher-stakes uses. Data governance covers provenance and documentation. Lifecycle monitoring tracks performance after deployment. Human oversight keeps accountable people in the loop. This section is informational and is not legal or regulatory advice.


The Commercial Landscape: How Organizations Use AI Biology Platforms


Organizations use AI in biology through internal teams, cloud and laboratory infrastructure, packaged platforms, open models and partnerships with AI-native biotechs.


  • Internal AI and machine-learning teams that build models on proprietary data.

  • Cloud and laboratory infrastructure for computing, storage and automation.

  • Protein-modeling and design platforms, and small-molecule discovery systems.

  • Genomics analytics and pathology or imaging platforms.

  • Open biological foundation models. The AlphaGenome authors, for example, provide tools for making predictions from sequence [6].

  • Partnerships with AI-native biotechs, where the partner brings models and the sponsor brings targets, data and development capacity.


Prices vary widely by vendor and contract, and public price data are scarce, so this guide does not quote them. Treat any price claim as something to confirm in a written quote.


Build vs Buy vs Partner: How Should a Biotech or Pharma Team Decide?


Build when the problem is core to your advantage and you hold unique data and talent. Buy when a validated tool solves a standard problem. Partner when you need capability you cannot hire quickly and are willing to share risk and control.


Approach

Best when

Strengths and weaknesses

Requirements and risks

Build

You own unique data and want a long-term platform

Full customization and data control; slow, costly and talent-dependent

ML engineers, biologists, computing and MLOps; risk of maintenance burden and unvalidated models

Buy

The task is standard and has public benchmarks

Fast start and vendor support; less customization and possible lock-in

Integration, security review and clear data and IP terms; risk of opaque methods

Partner

You need models plus expertise and want shared risk

Access to specialist teams; shared upside and shared control

Clear IP, data and governance agreements; risk of misaligned incentives


Across all three, decide who owns the data, who can fine-tune the model, how validation will work, how the tool fits your workflow and security rules, who pays for computing, and who maintains the system. Test lock-in by asking whether you can export your data, models and results.


How to Evaluate an AI Biology Platform or Vendor


Evaluate a vendor by asking what exact problem it solves, what evidence supports its performance against a meaningful baseline, and whether results were validated on data like yours. Table 5 is a compact checklist.


Evidence question

Why it matters

Red flag

What exact problem does it solve?

Defines success

Vague claims of a platform for everything

Which baseline was used?

Shows real gain

No comparison with simple methods [9]

Was it independently benchmarked?

Reduces bias

Only vendor-run tests

Was performance shown on new, out-of-distribution data?

Tests generalization

Random splits only

Is there experimental validation?

Separates prediction from reality

Retrospective results presented as proof

What training data and rights apply?

Legal and bias risk

Unclear licenses or provenance

Can outputs be interpreted, with uncertainty?

Supports decisions

Single scores without confidence

What integration and security are needed?

Drives real cost

No security documentation

What is the regulatory context?

Fits intended use

Ignores context of use [20]

Which KPI defines success?

Makes ROI measurable

No agreed metric


Also ask whether you need a foundation model at all. If a simpler model or standard tool meets your baseline, it may be cheaper and easier to validate. And ask whether you can reproduce the vendor's results on your own data.


How to Implement AI in a Biology or Drug-Discovery Organization


Start with one narrow problem, prove value against a baseline, and scale only when the evidence supports it.


  1. Define a narrow, high-value problem and the decision it informs.

  2. Establish baseline performance with current or simple methods.

  3. Audit the data: quality, labels, rights and bias.

  4. Choose build, buy or partner.

  5. Run a controlled pilot with domain scientists involved.

  6. Validate computationally, including on new data.

  7. Validate experimentally, or clinically where relevant.

  8. Measure operational and scientific outcomes.

  9. Document governance: owners, monitoring and change control.

  10. Scale only after the evidence supports it.


Measure return on investment as observed outcomes, not assumed percentages. Useful measures include decision quality, experimental hit rate, cycle time per design–test loop, cost per validated candidate and, in clinical settings, patient outcomes. Compare each against the baseline from step two.


Where AI in Biology Is Heading Next


Likely directions are multimodal models, richer perturbation data, causal modeling, AI agents and closed-loop laboratories. These are forecasts and emerging directions, not established results.


  • Multimodal models that combine sequence, structure, images and text, as in virtual-cell proposals [11].

  • Richer perturbation data to train and test models that must beat simple baselines [9].

  • Causal models that learn from interventions, not only observations.

  • AI agents and closed-loop laboratories, where models propose experiments and automated equipment runs them, as early work such as the Virtual Lab hints [18].

  • Model–experiment feedback loops that measure prediction errors and retrain.

  • Governance and biosecurity screening that keep pace with design tools [19].


For a broader and more speculative view, see our look at whether AI can solve the big open problems in biology.


So, How Is AI in Biology Transforming Research, Healthcare, and Drug Discovery?


AI in biology is transforming research by predicting structures, ranking variants, reading images and shortening the path from question to experiment. It is beginning to transform drug discovery, with AI-designed molecules in human trials and one program in Phase III by its developer's report [14]. It has not yet transformed healthcare outcomes at scale, and independent benchmarks show that simple baselines still matter [9][10]. The practical answer is to use AI where it beats a strong baseline, validate at the level of the claim, and let experiments and clinical trials decide.


Frequently Asked Questions


What is AI in biology?


AI in biology uses machine learning to find patterns in biological data and to predict or design proteins, cells, molecules and experiments. It supports research in genomics, imaging and drug discovery, and it works best alongside laboratory validation.


How is AI used in biological research?


Researchers use it to predict protein structures [1][2], rank DNA variants [5][6], classify cells and images [8][12], and prioritize experiments. The outputs are hypotheses that still need testing.


How is AI used in drug discovery?


It supports target identification, virtual screening, molecule generation, property prediction, biomarker work and trial planning. Rentosertib shows a generative-AI program reaching Phase III by its developer's report [14], but most uses sit earlier in the pipeline.


Does AI make drug development faster?


It can speed early computational steps, and AI-discovered molecules showed high Phase I success in one analysis [16]. Phase II success was comparable to historic averages on a small sample [16], so overall speed-ups are not yet demonstrated. Clinical trials, safety review and manufacturing remain.


What is a biological foundation model?


It is a large model pretrained on broad biological data, such as single-cell profiles, DNA or pathology images, and then adapted to specific tasks [7][8][12]. Its value is tested by comparison with simple baselines [9].


What is the difference between computational biology and AI?


Computational biology is the wider field of using computation to study biology. AI is one toolset within it, alongside statistics and simulation.


How does AlphaFold affect drug discovery?


AlphaFold gives structural hypotheses for targets, and AlphaFold 3 models complexes with ligands and nucleic acids [2]. It does not by itself predict binding strength, safety or efficacy.


Can AI design proteins?


Yes, with experimental support in some cases. RFdiffusion designs were tested in the laboratory, including a binder whose cryo-EM structure matched the design [4]. Success rates vary, so designs still need assays.


Can AI predict diseases from DNA?


AI can estimate the effects of some DNA variants, such as missense variants with AlphaMissense [5] and regulatory variants with AlphaGenome [6]. These are research and prioritization tools, not stand-alone diagnoses.


What are virtual cells?


Virtual cells are proposed AI models that simulate how cells behave under different conditions [11]. They are a research goal, and current models have not consistently beaten simple baselines at perturbation prediction [9].


Will AI replace biologists?


Current evidence points to collaboration. Even the Virtual Lab relied on human researchers for feedback and laboratory testing [18]. Roles may shift toward experiment design, data quality and validation.


What are the limitations of AI in biology?


Key limits include biased or noisy data, batch effects, weak generalization, benchmark leakage, limited interpretability, and the need for wet-lab and clinical validation.


How should companies evaluate AI drug-discovery platforms?


Define the problem, demand comparison with a simple baseline, ask for independent and out-of-distribution results, check data rights, and agree on a measurable KPI before a pilot.


What regulations apply to AI in drug development?


The FDA's January 2025 draft guidance proposes a risk-based credibility framework [20]. The FDA and EMA issued ten principles on 2026-01-14 [22], and EMA published a reflection paper in 2024 [24]. This is not legal advice.


What will AI in biology look like over the next few years?


Forecasts point to multimodal models, richer perturbation data, agents and closed-loop laboratories [11][18]. These are directions, not results, and progress will depend on data, benchmarks and validation.


Key Takeaways


  • Match the evidence level to the claim, because computational, experimental, clinical and regulatory evidence are different things.

  • Structure prediction is mature, while protein design, perturbation prediction and virtual cells are earlier-stage.

  • A simple baseline is a required test, not an insult to a new model [9][10].

  • High Phase I success shows that AI can design drug-like molecules, not that those drugs work better in patients [16].

  • A clinical candidate is not an approved medicine, and rentosertib remains investigational [14].

  • Most authorized clinical AI devices are in radiology, so most biology models still serve research [26].

  • Governance is part of the product: context of use, data rights and lifecycle monitoring [22].

  • The strongest teams pair models with fast experimental feedback.


Actionable Next Steps


  1. Write one sentence naming the biological question and the decision it will change.

  2. Inventory your data: quality, labels, permissions and bias.

  3. Set a simple baseline before testing any AI system.

  4. Ask vendors for independent, out-of-distribution and experimental evidence, using the checklist above.

  5. Run a small pilot with domain scientists and success metrics agreed in advance.

  6. Validate in the laboratory or clinic at the level the claim needs.

  7. Track hit rate, cycle time and cost per validated candidate against your baseline.

  8. Document governance, including context of use, owners, monitoring and human oversight, and check the current FDA and EMA pages before regulated use.


Glossary


  • ADMET: Absorption, distribution, metabolism, excretion and toxicity of a drug candidate.

  • Artificial intelligence: Computer systems that perform tasks that seem to need intelligence.

  • Baseline: A simple or current method used as a fair point of comparison.

  • Biological foundation model: A large model pretrained on broad biological data and adapted to many tasks.

  • Computational biology: Use of computing, statistics and simulation to study living systems.

  • Context of use: The exact role and scope of a model in a decision.

  • Deep learning: Machine learning with many-layered neural networks.

  • Diffusion model: A generative model that learns to remove noise step by step.

  • Embedding: A compact numerical summary of an item such as a cell or sequence.

  • Foundation model: A large model pretrained on broad data, then adapted.

  • Generative AI: Models that create new data, such as sequences, structures or images.

  • Genomic language model: A language-style model trained on DNA sequence.

  • Machine learning: Methods that learn patterns from examples.

  • Multimodal model: A model that combines data types such as text, images and omics.

  • Omics: Large-scale measurements such as genomics, transcriptomics and proteomics.

  • Perturbation: A deliberate change to a cell, such as switching off a gene.

  • Protein structure prediction: Estimating a protein's 3D shape from its sequence.

  • Single-cell sequencing: Measuring gene activity in individual cells.

  • Spatial biology: Measuring molecules while keeping their position in tissue.

  • Transformer: A neural network design that models relationships across a sequence.

  • Virtual cell: A proposed AI model that simulates cell behavior.


Sources & References


  1. Jumper J, et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (published online 2021-07-15). https://doi.org/10.1038/s41586-021-03819-2

  2. Abramson J, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493–500 (2024-06-13; online 2024-05-08). https://doi.org/10.1038/s41586-024-07487-w

  3. Royal Swedish Academy of Sciences. The Nobel Prize in Chemistry 2024, press release. 2024-10-09. https://www.nobelprize.org/prizes/chemistry/2024/press-release/

  4. Watson JL, et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023). https://doi.org/10.1038/s41586-023-06415-8

  5. Cheng J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science 381(6664), eadg7492 (2023-09-22). https://doi.org/10.1126/science.adg7492

  6. Avsec Ž, et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature 649, 1206–1218 (2026-01-28). https://doi.org/10.1038/s41586-025-10014-0

  7. Brixi G, et al. Genome modeling and design across all domains of life with Evo 2. Arc Institute manuscript; published in Nature, 2026. https://arcinstitute.org/manuscripts/Evo2.pdf

  8. Theodoris CV, et al. Transfer learning enables predictions in network biology. Nature 618, 616–624 (2023-05-31). https://doi.org/10.1038/s41586-023-06139-9

  9. Ahlmann-Eltze C, Huber W, Anders S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nature Methods 22(8), 1657–1661 (2025-08-04). https://doi.org/10.1038/s41592-025-02772-6

  10. Kedzierska KZ, Crawford L, Amini AP, Lu AX. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology (2025). https://doi.org/10.1186/s13059-025-03574-x

  11. Bunne C, et al. How to build the virtual cell with artificial intelligence: priorities and opportunities. Cell 187(25), 7045–7063 (2024-12-12). https://doi.org/10.1016/j.cell.2024.11.015

  12. Chen RJ, et al. Towards a general-purpose foundation model for computational pathology. Nature Medicine 30(3), 850–862 (2024-03-19). https://doi.org/10.1038/s41591-024-02857-3

  13. Xu Z, et al. A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial. Nature Medicine 31(8), 2602–2610 (online 2025-06-03). https://doi.org/10.1038/s41591-025-03743-2

  14. Insilico Medicine. Insilico Medicine doses first patient in GENESIS-IPF-3, the world's first Phase III trial of a generative AI-driven innovative drug (company press release). PR Newswire, 2026-09-09. https://www.prnewswire.com/news-releases/insilico-medicine-doses-first-patient-in-genesis-ipf-3-the-worlds-first-phase-iii-trial-of-a-generative-ai-driven-innovative-drug-302873749.html

  15. Buntz B. Insilico's AI-designed rentosertib shows promise in first phase 2a trial results. Drug Discovery and Development, 2025-06-05. https://www.drugdiscoverytrends.com/insilicos-ai-designed-rentosertib-shows-promise-in-first-phase-2a-trial-results/

  16. Jayatunga MKP, et al. How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons. Drug Discovery Today 29(6), 104009 (2024). https://doi.org/10.1016/j.drudis.2024.104009

  17. BenevolentAI. Top-line Phase IIa results for BEN-2293 in mild-to-moderate atopic dermatitis (company press release). Business Wire, 2023-04-04. https://www.businesswire.com/news/home/20230404006090/en

  18. Swanson K, Wu W, Bulaong NL, Pak JE, Zou J. The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies. Nature (2025-07). https://doi.org/10.1038/s41586-025-09442-9

  19. Wittmann BJ, et al. Strengthening nucleic acid biosecurity screening against generative protein design tools. Science 390(6768), 82–87 (2025-10). https://doi.org/10.1126/science.adu8578

  20. U.S. Food and Drug Administration. Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance for industry. January 2025. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-artificial-intelligence-support-regulatory-decision-making-drug-and-biological

  21. U.S. Food and Drug Administration. Federal Register notice of availability of the AI draft guidance. 2025-01-07. https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug

  22. U.S. FDA and European Medicines Agency. Guiding Principles of Good AI Practice in Drug Development. January 2026. https://www.fda.gov/media/189581/download

  23. U.S. Food and Drug Administration. Guiding Principles of Good AI Practice in Drug Development, web page (2026-01-14). https://www.fda.gov/about-fda/artificial-intelligence-drug-development/guiding-principles-good-ai-practice-drug-development

  24. European Medicines Agency. Reflection paper on the use of artificial intelligence in the medicinal product lifecycle. Adopted September 2024. https://www.ema.europa.eu/en/scientific-guidelines/use-artificial-intelligence-ai-medicinal-product-lifecycle

  25. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. 2023-01-26. https://doi.org/10.6028/NIST.AI.100-1

  26. Golshani P, Joseph M. Three decades of FDA authorizations of AI/ML-enabled medical devices: persistent specialty concentration and the care-delivery gap (1995–2025). medRxiv preprint (2026). https://doi.org/10.64898/2026.05.08.26352766


Company press releases [14][17] are cited only for company-reported facts, and the medRxiv preprint [26] is not peer reviewed.

 
 
bottom of page