top of page

Can AI Solve the 12 Millennium Problems for Biology?

6 hours ago
25 min read
AI visualizing 12 Millennium Problems for Biology.

On September 19, 2026, FutureHouse and Edison Scientific first published a deliberately hard test for biology: twelve challenges that a laboratory can verify, from a living cell that reads four-letter codons to a mouse that regrows a limb. AI can already design binders, read thousands of papers in minutes and analyse assay data, yet in the sources checked for this article, dated September 29, 2026, no verified report shows any of the twelve meeting its official criterion. The real question is whether AI will solve these problems itself, or mainly speed up the scientists, robots and laboratories that do.


TL;DR


  • The Millennium Problems for Biology are twelve living, laboratory-testable challenges assembled by Sam Rodriques and Michaela Hinks for FutureHouse and Edison Scientific. They are a new proposal, not a ratified community list, and no prize amounts have been verified.

  • As of September 29, 2026, no verified result satisfies any official criterion. Adjacent advances such as BindCraft, DISCO, reverse-translation peptide sequencing and FGF2/BMP2 digit regeneration are meaningful progress, but each falls short.

  • Protein-design problems (Rubisco, proteases, binders, nitrogenases, 5′ polymerases) are the most computationally tractable. Whole-organism problems (cryopreservation, limb regeneration) are dominated by physiology, not model quality.

  • Every official criterion demands preregistered, blinded wet-lab proof, so the decisive evidence will be experimental, whatever AI contributes.

  • The most realistic path is capable AI embedded in automated design-build-test-learn loops, not a chatbot solving biology alone.


The Short Answer


Probably not alone. AI can already design proteins, synthesise literature and analyse assay data, but the twelve Millennium Problems for Biology require preregistered, blinded experiments, and none has been verified as solved. The likeliest outcome is AI sharply accelerating human-run, increasingly automated laboratories, with protein-design problems moving first.

What is the biggest barrier to AI solving biology’s 12 Millennium Problems?

  • 0%Wet-lab validation

  • 0%Missing biological data

  • 0%Model reliability and reasoning

  • 0%Whole-organism complexity


Table of Contents



What Are the 12 Millennium Problems for Biology?


The Millennium Problems for Biology are twelve open problems published at millenniumproblems.bio by Edison Scientific and FutureHouse. The About page records Version 1.0 as published on 2026-09-19 with a last update on 2026-09-21, and FutureHouse’s announcement followed on September 23, 2026. Sam Rodriques and Michaela Hinks developed the initial list, and the project is described as living, so criteria may change.


The design rule is that each problem should be very hard to solve but easy to validate in a simple laboratory. Broad goals such as “Solve Aging” were excluded because they are hard to validate. The organisers call the set, in some sense, the “last reasonable eval” for AI in biology.


This is a newly proposed catalogue, not an equivalent of the Clay Mathematics Institute program. CMI announced its problems in 2000 and attached $1 million to each. No comparable biology prize has been established in the sources checked. A post attributed to Rodriques on X says FutureHouse will convene a panel to adjudicate success and announce cash prizes, but gives no amounts.


The table below condenses each challenge. Exact criteria follow in the numbered sections.


Problem

Official goal in plain English

Main bottleneck

Where AI helps

1. Origins of life

Self-replicating protocells from plausible chemistry

Chemistry of heredity plus compartments

Search over ribozyme and reaction space

2. Cryopreservation

Freeze and revive whole adult mice, >99% viability

Whole-body ice and thermal damage

Modelling protocols; limited

3. Reverse translatase

Enzyme that reads peptides into nucleic acid

No natural template-free mechanism

Enzyme design and screening

4. Improve Rubisco

Beat natural specificity and speed together

Catalytic trade-offs, slow assays

Sequence models, active learning

5. Quadruplet cell

Living cell using only four-base codons

Rebuilding all translation machinery

RNA and protein design

6. Limb regeneration

Regrow a functional adult mouse limb

Tissue patterning and innervation

Hypothesis and target search

7. Bacterial gene therapies

Infectious AAV and lentivirus made in bacteria

Assembly, genome packaging, infectivity

Capsid and factor design

8. Programmable proteases

Design proteases for any blinded site

Specificity in living cells

Generative enzyme design

9. Cell-penetrating binders

Extracellular binders that hit intracellular targets

Endosomal escape

Binder design

10. Protein amplification

Copy peptides exponentially without nucleic acid

No known mechanism

Speculative mechanism search

11. 5′ polymerases

Four 3′→5′ polymerases matching Taq-class tools

Fidelity and processivity

Enzyme redesign

12. New nitrogenases

Non-homologous enzyme that fixes N₂

Metal cofactor assembly

De novo design, cofactor modelling


Why These Problems Are an Unusual Test for AI


They are unusual because success is defined by preregistered laboratory outcomes, not by model scores. Several criteria require 20 or 100 preregistered targets, a fixed success rate, designs produced within 24 hours and no target-specific wet-lab work once targets are revealed. That structure is built to defeat cherry-picking.


The mathematics comparison is instructive. CMI’s rules require a proposed solution to be published in a qualifying outlet, to be at least two years old and to have received general acceptance before CMI considers it. On September 11, 2026, CMI’s Navier–Stokes announcement said the problem had “apparently been settled” rather than declaring a prize awarded. A claim, expert scrutiny and official recognition are different stages, and biology needs the same distinctions.


Biology adds a second difficulty: validation is physical. A model can rank a million candidates in seconds, but a cell, an enzyme or a limb must actually work.


How AI Is Already Changing Biological Discovery


AI has changed the front end of discovery most: structure prediction, sequence-to-function models, generative design and agents that read and analyse. AlphaFold 3 (Nature, 2024) predicts joint structures of biomolecular complexes. AlphaGenome (Nature, January 28, 2026) predicts thousands of genomic signals from up to one million base pairs and matched or exceeded leading external models in 25 of 26 variant-effect evaluations.


Generative design followed. RFdiffusion (Nature, 2023) and ProteinMPNN (Science, 2022) made backbone and sequence design routine. BindCraft (Nature, 2025) reported experimental binder success rates of 10–100% across targets, averaging 46.3%, without high-throughput screening.


Agents are newer. Robin (Nature, May 19, 2026) generated hypotheses, proposed assays and analysed data for dry age-related macular degeneration, but the authors state that human scientists ran the laboratory experiments. Its top compound, ripasudil, raised retinal pigment epithelium phagocytosis 1.89-fold over control in the agent’s analysis and 1.75-fold in the human reanalysis, in vitro only. Kosmos, an Edison Scientific preprint, reported 79.4% of sampled report statements accurate, but only 57.9% of synthesis statements.


These advances are strong inputs to the twelve problems. None is a solution to one.


1. Origins of Life


The official challenge asks for the unassisted emergence of self-replicating RNA- and protein-based cells from a plausible primordial soup and energy source. Emergent cells must increase in abundance at least 10⁶-fold (roughly 20 generations), plausibly keep dividing indefinitely, and carry clearly heritable information that causally helps determine their composition. Solving it would show how chemistry becomes biology.


The strongest recent work concerns the replicator, not the cell. A 2026 Science paper describes QT45, a 45-nucleotide polymerase ribozyme found in random sequence pools that is reported to synthesise its complementary strand and itself from trinucleotide substrates in eutectic ice. A 2025 Nature Chemistry study showed RNA replication cycles driven by pH and freeze–thaw. Synthetic protocells that support DNA self-replication with protein synthesis exist, but they are built from purified, engineered parts. This is meaningful progress, not completion of the official challenge.


AI can search ribozyme sequence space, as generative models of self-reproducing ribozymes did in 2025, and prioritise prebiotic reaction networks. It cannot bypass the missing step: unassisted emergence under plausible conditions without hand-fed reagents or engineered machinery. The decisive bottleneck is a working chemical route that links replicators to compartments that divide and inherit. A convincing result would show 10⁶-fold expansion with unambiguous heredity.


2. Cryopreservation


The challenge asks for reversible whole-body freezing or vitrification of live, intact, wild-type adult mice for at least 24 hours, with recovery at more than 99% viability, no permanent organ damage and ethics approval. Vitrification means cooling to a glass-like state without ice crystals. Success would matter for transplant logistics, critical care and any medicine that needs time.


Organ-scale results are real but different. Han and colleagues (Nature Communications, 2023) vitrified rat kidneys, rewarmed them by nanowarming (magnetic heating of iron-oxide nanoparticles) and achieved life-sustaining transplantation. A commentary describes storage for up to 100 days. A 2025 liter-scale study pushed physical vitrification and nanowarming toward human-organ volumes. These are single perfused organs, not whole living animals.


AI can help model cryoprotectant mixtures, thermal stress and warming rates, but data are sparse and the physics spans scales. It cannot bypass whole-body toxicity, uneven cooling, heart and brain injury, or the need for animal experiments. The decisive bottleneck is whole-organism damage during cooling, cryoprotectant exposure and rewarming, not prediction. A convincing result is 24 hours in a frozen or vitrified state and recovery of more than 99% of animals with normal organ function, independently replicated.


3. The Reverse Translatase


The challenge asks for a purified protein catalyst or fixed complex that processively reads an untagged polypeptide and writes a covalent nucleic-acid encoding of its residue sequence. It may use no nucleic-acid template, preattached barcode, residue-specific operator cycle or database lookup. Testing requires at least 100 preregistered random peptides of 50 or more residues, at least 90% sequence accuracy, an average read length of at least 25 residues and average quality of at least Q10. Success would turn protein sequencing into DNA sequencing.


A 2026 Nature Biotechnology paper uses “reverse translation” for single-molecule peptide sequencing, and a News & Views piece covered it. The method applies a modified Edman degradation that iteratively releases N-terminal amino acids tagged with peptide-specific DNA barcodes, then reads the barcodes by DNA sequencing. It is a strong chemistry-and-sequencing workflow, but it relies on barcodes and iterative chemical cycles, the very mechanisms the official criterion excludes. It is not a reverse-translatase enzyme.


AI can design and triage candidate enzymes, but nature offers no template-free starting point. The decisive bottleneck is mechanism: one catalyst must recognise twenty amino acids in sequence context and polymerise nucleotides accordingly. A convincing solution passes the preregistered 100-peptide test at the stated accuracy, read length and quality.


4. Improve Rubisco


Rubisco fixes carbon dioxide in photosynthesis, but it is slow and also reacts with oxygen. The challenge asks for one enzyme whose CO₂/O₂ specificity (Sc/o) is at least that of Galdieria partita Rubisco and whose turnover number (kcat) is at least that of maize Rubisco, measured in paired assays with both controls. The enzyme may be designed, discovered or evolved. Success would beat the natural trade-off frontier and could lift crop and bioproduction efficiency.


Progress is incremental. A July 2026 bioRxiv preprint on machine-learning-assisted directed evolution of plant Rubisco reports variants with improved catalytic efficiency, including single substitutions that raise a catalytic parameter without compromising, and sometimes improving, specificity. A 2026 AI Chemistry paper shows that active learning on protein language model embeddings recovers top Rubisco mutants more efficiently than random sampling at a fixed experimental budget. Neither is reported to beat both official controls at once.


AI fits well here: language models, structure-informed design and active learning reduce the variants that must be built. It cannot bypass careful kinetic measurement, and folding and assembly in host cells remain difficult. The decisive bottleneck is assay throughput plus the apparent trade-off between speed and specificity. A convincing result is one enzyme that meets both thresholds in paired assays, independently replicated.


5. The Quadruplet Cell


Nearly all life reads genes in three-base codons. The challenge, contributed by Erika Alden DeBenedictis, asks for a living, replicating cell in which every protein-coding sequence, including the translation machinery itself, is written in uninterrupted, non-overlapping quadruplet codons, with no detectable triplet decoding. The code must be genuinely four-base: a mutation should be about equally likely to change the amino acid at any codon position, so a code where the first three bases nearly always suffice does not count. Success would create a genetically isolated organism.


Quadruplet decoding already exists inside otherwise triplet cells. DeBenedictis, Söll and Esvelt showed that E. coli tRNAs can be evolved to read four-base codons, and a review notes that all twenty E. coli tRNA isoacceptor classes can be converted into quadruplet suppressors, nine of them forming a mutually orthogonal set. A Chemical Reviews article describes orthogonal ribosomes that translate quadruplet codons. These add a few quadruplet codons beside triplet decoding. They are expansions, not an all-quadruplet cell.


AI could design tRNAs, synthetases and ribosomal RNA variants and predict the fitness of recoded genes. It cannot bypass the scale of the rewrite: every gene and the whole translation apparatus must change together in a living cell that still carries triplet decoders. The decisive bottleneck is genome-wide recoding and eliminating triplet decoding. A convincing result is a replicating cell with verified quadruplet-only translation that passes the mutation-symmetry test.


6. Somatic Limb Regeneration


The challenge asks for reproducible regrowth of an amputated limb in an adult wild-type mouse. The regrown limb must perform like controls in a standard motor battery and in sensory perception, and blinded observers must be unable to tell which limb regrew from non-invasive observation, all under ethics approval. Success would open a route to regenerative treatment after injury.


The best recent evidence concerns digits. A Nature Communications paper (April 17, 2026) reports that sequential FGF2 then BMP2 treatment of a non-regenerating mouse digit amputation replaced wound fibrosis with regeneration of phalangeal and sesamoid bones, tendon and ligament, a synovial joint and articular cartilage. That is a genuine advance in mammalian skeletal regeneration. It is a digit, not a limb, and the reported outcomes are structural rather than the official motor, sensory and blinded-observer tests.


AI can help prioritise signalling factors, model tissue patterning and score wound-healing programmes. It cannot bypass the biology: a limb needs bone, muscle, nerve, vessel and skin to pattern and reconnect together, while mammalian wounds favour scarring. The decisive bottleneck is coordinated patterning and innervation over long animal experiments. A convincing result is a reproducible regrown adult mouse limb that passes the motor, sensory and blinded tests across cohorts.


7. Bacterial Production of Gene Therapies


Adeno-associated virus (AAV) and lentivirus are workhorse gene-therapy vectors, normally made in mammalian or insect cells. The challenge asks for infectious, replication-incompetent AAV and lentivirus produced in bacteria, carrying a prespecified viral genome. Capsid-to-genome and infectious-unit-to-genome ratios must resemble mammalian production, and genomes must be nuclease-resistant. AAV alone counts as partial success. Success could lower manufacturing cost and complexity.


Current bacterial work makes empty shells. A 2024 Protein Expression and Purification study assembled empty VP3-only AAV2 virus-like particles inside E. coli at low yield. A 2019 Scientific Reports paper assembled AAV2 VP3 capsids from E. coli protein in a chemically defined reaction, and a synthetic biology study formed AAV5 virus-like particles in E. coli. A virus-like particle is a capsid without the required genome and without demonstrated infectivity, so none of these meets the criterion.


AI could redesign capsid proteins and assembly factors for bacterial folding and predict packaging interactions. It cannot bypass the missing capabilities: genome packaging, the full capsid protein set and infectivity. The site itself anticipates that lentivirus will be much harder. The decisive bottleneck is packaging a specified genome into functional particles in bacteria. A convincing result shows nuclease-resistant genomes and infectious-unit ratios comparable to mammalian production, validated independently under biosafety oversight.


8. Programmable Proteases


A protease is an enzyme that cuts proteins. The challenge asks for proteases designed on demand for a blinded, accessible site in an endogenous folded protein, cutting efficiently in living cells. Catalytic efficiency and proteome-wide off-target cleavage must match or beat widely used site-specific proteases. The test uses 20 preregistered sites with more than 80% success, designs within 24 hours and no screening or target-specific evolution after sites are revealed. A weaker peptide-in-solution version counts as partial success.


Recent work is early and largely preprint. A bioRxiv preprint (Proteus2, posted January 2026, revised March 2026) designed zinc metalloproteases for three cleavage sites in amyloid-β peptide and reports five validated enzymes with high specificity, while stating that de novo design had not yet produced proteases that cleave any desired bond in a native protein with high precision. CleaveNet (Nature Communications, January 2026) designs substrates for a known protease, which is the reverse problem. Peptide cleavage in vitro is at most the weaker partial form.


Generative enzyme design, structure prediction and active learning all apply here. AI cannot bypass folded-protein context: sites must be accessible, cleavage must be specific across the proteome and cells add competing proteins. The decisive bottleneck is specificity and activity on folded targets in cells without screening. A convincing result is success above 80% across 20 blinded sites, with proteome-wide off-target profiling.


9. Cell-Penetrating Protein Binders


Most biologic drugs, including antibodies, act outside cells. The challenge asks for zero-shot protein binders that reliably engage preregistered intracellular targets in living cells when administered extracellularly at pharmacologically supported concentrations, without further optimisation. Entry must be therapeutically plausible, so transfection, intrabody expression, electroporation and membrane disruption are excluded. The test uses 20 targets, an 80% success rate, designs within 24 hours and no target-specific wet-lab work after disclosure. Success would open many currently undruggable targets.


Design is the strong half. BindCraft reported nanomolar binders with experimental success rates of 10–100% across targets that included cell-surface receptors, allergens and CRISPR-Cas9, but those experiments did not test extracellular delivery to intracellular targets in living cells. Delivery is the weak half. A review of antibody delivery calls endosomal escape the major bottleneck, and a February 2026 preprint on an engineered membrane translocation domain reports cytosolic delivery at low nanomolar concentrations while noting that post-escape aggregates are a recognised bottleneck.


AI can design the binder, co-design a cell-entry module and model membrane translocation. It cannot bypass endosomal trapping, serum stability, toxicity or the need for cell assays across 20 targets. The decisive bottleneck is delivery, not binding. A convincing result is a therapeutically plausible extracellular entry mechanism that engages targets at an 80% success rate across 20 preregistered targets, with no optimisation after disclosure.


10. Protein Amplification Chain Reaction


The polymerase chain reaction amplifies DNA exponentially, but no equivalent exists for proteins. The challenge asks for input-protein-dependent synthesis of new, full-length, sequence-faithful covalent polypeptide copies from amino-acid monomers, without a nucleic-acid template or preformed cognate scaffold, in a single pot. Testing needs at least 100 preregistered random peptides of 50 or more residues, at least 1,000-fold amplification and at least 90% per-residue accuracy. Explicit peptide sequencing and reverse translation to a nucleic-acid intermediate are excluded.


No verified work meets this bar. Biology copies sequence information through nucleic acids, and no peer-reviewed article or preprint reporting arbitrary-sequence peptide copying without a nucleic-acid template turned up in the sources checked for this article. That makes this the most open-ended problem on the list, and the one where the state of the art is least defined.


AI’s role is speculative: searching mechanism space for catalysts that recognise a peptide and template assembly of a copy from monomers, and proposing candidate scaffolds. It cannot bypass the absence of a known mechanism or the fidelity demanded of every residue. The decisive bottleneck is conceptual, not computational. A convincing result is a single-pot reaction that passes the preregistered 100-peptide test without sequencing or nucleic-acid steps.


11. 5′ Polymerases


All known DNA polymerases, RNA polymerases, reverse transcriptases and RNA-dependent RNA polymerases (RdRPs) extend nucleic acids in the 5′→3′ direction. The challenge asks for a complete set of 3′→5′ polymerases: a DNA polymerase, an RNA polymerase, a reverse transcriptase and an RdRP, with processivity and error characteristics similar to or better than Taq, T7 RNA polymerase, M-MLV reverse transcriptase and Phi6 RdRP. They may be designed, discovered or engineered. Success would give molecular biology a mirror-image set of tools.


Nature offers a foothold, not a toolkit. The Thg1 and Thg1-like protein (TLP) family catalyses templated synthesis of RNA in the reverse direction, and a review of the Thg1 superfamily describes its 3′-to-5′ polymerization. Structural work on a TLP shows how template-dependent elongation proceeds in that direction. These enzymes have been characterised largely on tRNA-related substrates and short RNA extensions. The criterion requires four enzyme classes, including DNA synthesis and reverse transcription, at Taq-class performance.


AI can redesign polymerase active sites, predict fidelity determinants and guide directed evolution. It cannot bypass the chemistry: 3′→5′ extension requires an activated 5′ end on the growing strand, and Thg1 uses a multi-step activation reaction (adenylylation, nucleotidyl transfer, pyrophosphate removal) unlike canonical polymerases. The decisive bottleneck is fidelity and processivity across four enzyme classes. A convincing result is all four enzymes benchmarked against Taq, T7 RNA polymerase, M-MLV and Phi6 in matched assays.


12. New Nitrogenases


Nitrogenase converts atmospheric nitrogen (N₂) to ammonia, the basis of biological nitrogen fixation. The challenge asks for a protein that reduces N₂ to ammonia at rates of a similar order of magnitude to natural nitrogenases while falling well below sequence- and structure-similarity thresholds to all known nitrogenase and nitrogenase-like proteins. It may be designed, discovered or evolved. Success would show that a complex metalloenzyme can be invented, with implications for fertiliser and agriculture.


Current work reconstructs or studies natural enzymes. A January 2026 Nature Communications study resurrected ancestral nitrogenases in a bacterial host to test isotope signatures across more than two billion years. An eLife study combined ancestral reconstruction, crystallography and deep-learning structure prediction to trace nitrogenase structural evolution. A PNAS study showed a Roseiflexus nitrogenase assembling its cofactor without the usual scaffold NifEN. All start from natural nitrogenases, so none is a non-homologous enzyme, and no verified AI-designed functional nitrogenase turned up in the sources checked.


AI can help design metal-binding sites, predict cofactor–protein interactions and search for distant natural relatives. It cannot bypass cofactor chemistry: the active-site metal clusters are assembled by dedicated proteins, and a well-folded design is inert without them. The decisive bottleneck is cofactor assembly and functional testing. A convincing result is measured ammonia production at natural-order rates from a protein below the similarity thresholds.


Comparing the 12 Problems: Where AI Helps and Where Biology Still Dominates


AI leverage is highest where the deliverable is a molecule that can be tested quickly, and lowest where the deliverable is a living animal or an undiscovered mechanism. The table compares the nature of each bottleneck. It deliberately avoids scores and probabilities, because the evidence does not support numerical forecasts.


Challenge

Computational leverage

Wet-lab dependence

Strongest enabling advance

Missing capability

1. Origins of life

Moderate: sequence and reaction search

Very high

QT45 polymerase ribozyme (Science, 2026)

Dividing, heritable protocell from plausible chemistry

2. Cryopreservation

Limited

Very high; whole animals

Rat kidney vitrification and nanowarming

Whole-body recovery above 99% viability

3. Reverse translatase

Moderate: enzyme design

High

Barcode-based reverse-translation sequencing

Template-free enzymatic mechanism

4. Improve Rubisco

High: fitness landscapes

Moderate; slow kinetics

ML-guided directed evolution (preprint)

Both thresholds in one enzyme

5. Quadruplet cell

Moderate: RNA and protein design

Very high; genome scale

Orthogonal quadruplet tRNAs and ribosomes

A cell with no triplet decoding

6. Limb regeneration

Limited

Very high; long animal studies

FGF2 then BMP2 digit regeneration

Functional, innervated limb

7. Bacterial gene therapies

Moderate: capsid design

High; biosafety-gated

Empty AAV particles in E. coli

Genome packaging and infectivity

8. Programmable proteases

High: generative design

High; cell assays

Proteus2 amyloid-β proteases (preprint)

General cleavage of folded proteins

9. Cell-penetrating binders

High for binding, low for entry

High

BindCraft one-shot binders

Endosomal escape at scale

10. Protein amplification

Speculative

High

None identified

A working mechanism

11. 5′ polymerases

Moderate: enzyme redesign

High

Thg1/TLP 3′→5′ RNA activity

Four Taq-class enzymes

12. New nitrogenases

Moderate: design plus cofactor modelling

High; cofactor assembly

Ancestral reconstruction with structure prediction

Non-homologous, active enzyme


The AI + Laboratory Stack These Problems Will Require


A design-build-test-learn (DBTL) loop needs five layers: models that propose designs, automated construction of DNA and proteins, high-throughput assays, data infrastructure that records provenance and failures, and an agent layer that chooses the next experiment, often through active learning. Robin shows the reasoning layer working, but its Nature paper says it does not yet produce precise, executable protocols, and its data-analysis agent Finch relies on expert prompts and scored 22.8% on a 170-question BixBench panel of bioinformatics and statistics tasks.


For protein-design problems the loop can be fast. For limb regeneration and cryopreservation each cycle needs animal facilities and weeks to months, so automation helps less. That asymmetry, more than model quality, will shape which challenges move first.


Commercial Implications: What Biotech Teams Should Build, Buy, or Partner For


Buy commoditising computational capability, build around proprietary data and assays, and partner for scarce physical throughput. Evaluate every vendor on experimentally validated hit rates, not benchmark scores; the difference is the theme of this whole list, and AI accuracy metrics rarely capture it.


What is available today


  • Literature and reasoning agents. Edison Scientific describes a platform of literature, analysis and molecular-design agents. Its May 19, 2026 announcement says the rearchitected Kosmos is available by application, core agents stay in the Edison Playground with 10 free credits a month, and Incyte is a named partner. Earlier pricing has changed, so confirm current terms. Grounding matters: in Robin’s ablation, an unaugmented model hallucinated 44.5% of references in assay proposals, versus none for the Crow literature agent.

  • Structure and genome models. AlphaFold 3 weights and outputs are free only for non-commercial use by non-commercial organisations under its terms, so commercial teams need other terms or alternatives. The AlphaGenome API is offered free for non-commercial use under separate terms.

  • Protein design. BindCraft is described as open source, and RFdiffusion, ProteinMPNN and DISCO (a preprint-stage enzyme co-design model) are published methods. As generative modelling tools open up, differentiation shifts to assays, data and validation.


A build, buy or partner framework


  • Strategic differentiation: build where the capability is the moat, buy where it is table stakes.

  • Data sensitivity: proprietary assay data favour internal or private deployments; check data-use terms.

  • Frequency of use: frequent workloads justify teams and infrastructure; occasional ones suit pay-per-use.

  • Model customisation: fine-tuning on internal data favours building or a partner that supports it.

  • Wet-lab integration: closed loops need assay and robot integration, so partner with a contract research or automation provider if capacity is missing.

  • Validation and regulatory burden: GxP settings need documented validation, audit trails and model version control.

  • Switching costs and talent: proprietary formats create lock-in, and people who span machine learning and experimental biology are scarce.

  • Total cost and reproducibility: count assay, compute and repeat-run costs, and demand prospective, blinded hit rates.


Data infrastructure sets the loop’s speed: standardised assay formats, machine-readable ELN and LIMS records, provenance, metadata and negative results. Robin’s Finch agent analysed raw flow cytometry and RNA-seq files, which only works when those records are complete.


What Would Count as a Real AI-Driven Solution?


A real AI-driven solution is one where AI originated the decisive design or insight and physical experiments then met the official criterion. Assistance alone does not qualify. The spectrum runs from weak to strong contribution:


  • Search and retrieval: an agent finds and summarises evidence.

  • Coding and data-analysis assistance.

  • Hypothesis generation: AI proposes a testable idea.

  • Molecule design: AI proposes sequences or structures.

  • Experiment selection and result analysis.

  • Updating models and hypotheses from new data.

  • A closed experimental loop run by an autonomous agent.

  • AI originates the decisive solution while humans execute the experiments.

  • Highly autonomous robotic science.


Attribution becomes difficult because a result may depend on foundation models, training datasets, software, instruments, technicians, contract research organisations, engineers and scientists. Robin’s authors state that all hypotheses, experimental directions, analyses and main-text figures were produced by the system, yet people ran the experiments and reviewed ranked candidates. “AI solved it” should be a claim about which steps AI performed, supported by logs, not a marketing shorthand.


Safety, Biosecurity, Ethics, and Reproducibility


Biosafety and biosecurity concerns rise with capability. Vector production, cell-entry proteins and enzymes that act on human proteins are dual-use topics, so responsible access, model safeguards and institutional review matter. Robin’s authors describe guardrails that include prioritising candidates with established safety profiles and screening platform queries with a classifier. Animal welfare and ethics approval are written into the cryopreservation and limb-regeneration criteria. For a broader framework, see this guide to AI risk management.


Reproducibility carries the rest. Preregistration, blinded evaluation and independent replication defend against benchmark gaming, publication bias and selective reporting. Data and model provenance, calibrated uncertainty and checks for hallucinated citations reduce error. This article describes mechanisms, criteria and bottlenecks; it deliberately gives no experimental protocols.


So, Can AI Solve the 12 Millennium Problems for Biology?


Not by itself, and not on the evidence available on September 29, 2026. AI can plausibly compress the hypothesis, design and search cycle for the molecular problems, but every official criterion is settled by physical, preregistered experiments.


Some challenges are unusually compatible with AI-driven molecular design: Rubisco, programmable proteases and the binder half of intracellular binders. Others are dominated by whole-organism physiology (cryopreservation, limb regeneration), by an unknown mechanism (protein amplification) or by physical constraints that no model removes: delivery, cofactor assembly, genome-scale recoding and biosafety-gated manufacturing.


The most realistic near-term model is increasingly capable AI embedded in automated design-build-test-learn systems operated by scientists, not a disembodied chatbot solving biology. If any problem falls, expect shared credit and a replicated experiment as proof. The list is living, so these conclusions apply to the criteria as published through September 21, 2026.


FAQ


What are the 12 Millennium Problems for Biology?


They are origins of life, cryopreservation, the reverse translatase, improve Rubisco, the quadruplet cell, somatic limb regeneration, bacterial production of gene therapies, programmable proteases, cell-penetrating protein binders, the protein amplification chain reaction, 5′ polymerases and new nitrogenases. Each has quantitative success criteria at millenniumproblems.bio.


Who created the Millennium Problems for Biology?


Sam Rodriques and Michaela Hinks developed the initial list with feedback from the FutureHouse and Edison Scientific teams. The official site records Version 1.0 as published on 2026-09-19, and FutureHouse maintains it as a living project.


Are they related to the Millennium Prize Problems in mathematics?


They draw inspiration from them but differ in origin, governance and validation. The Clay Mathematics Institute’s seven problems date from 2000, carry $1 million each, and require publication, two years and general acceptance before consideration. The biology list has no such institutional process yet.


Is there prize money for solving them?


No prize amount has been verified. A post attributed to Sam Rodriques says FutureHouse will convene a panel to adjudicate success and announce cash prizes. Check millenniumproblems.bio and FutureHouse for official updates.


Has AI solved any of the 12 biology problems yet?


No verified report in the sources checked, as of September 29, 2026, meets any official criterion. Related advances exist, such as BindCraft binders, DISCO enzyme designs (preprint) and FGF2/BMP2 digit regeneration, but each falls short of the exact challenge.


Which of the problems is most directly suited to protein-design AI?


Improve Rubisco, programmable proteases, the design half of intracellular binders, new nitrogenases and 5′ polymerases all involve designing or engineering proteins. Even so, assays, cell delivery, cofactor assembly and fidelity, not design alone, remain the limiting steps.


Can AlphaFold solve these problems?


No. AlphaFold 3 predicts structures of biomolecular complexes, but the challenges require functional molecules, cells or animals that pass preregistered experiments. Its weights and outputs are also free only for non-commercial use by non-commercial organisations.


What is a reverse translatase, and is reverse-translation peptide sequencing the same thing?


A reverse translatase would be a single enzyme that reads a peptide and writes nucleic acid encoding it. The 2026 Nature Biotechnology method uses DNA barcodes and iterative Edman chemistry, which the official criterion excludes, so it is meaningful progress but not the challenge.


Could AI help humans regenerate limbs?


It may help identify signals and model patterning. The 2026 mouse study regrew digit bones, joint tissue, tendon and ligament with FGF2 then BMP2. That is not a functional adult limb, and no verified route to human limb regeneration exists.


Why produce gene therapies in bacteria, and why is it hard?


Bacteria could offer simpler, cheaper manufacturing than mammalian cells. So far bacterial work has assembled empty AAV virus-like particles at low yield. Packaging a specified genome and producing infectious, nuclease-resistant vectors, and lentivirus especially, remain unshown.


Why are intracellular protein binders difficult?


Two problems must be solved together: designing a binder that engages the target, and delivering it from outside the cell into the cytosol at therapeutic concentrations. Endosomal escape is widely described as the major delivery bottleneck.


What would count as AI “solving” one of these challenges?


AI would have to originate the decisive design or insight, and physical experiments would then meet the official preregistered criterion. Because models, data, instruments and people all contribute, credible claims should document which steps the AI performed.


Key Takeaways


  • Read the criteria first: each problem defines success as a preregistered laboratory result, so adjacent progress is not completion.

  • The list is new, living and unratified. Watch the official site for changes, an adjudication panel and any prizes.

  • Design-centred problems can compress fastest, while whole-organism physiology moves slowest.

  • Delivery, cofactor assembly, genome packaging and genome-scale recoding are physical constraints that better models do not remove.

  • Robin shows literature-driven hypothesis generation and data analysis working, but humans ran its experiments, and Kosmos accuracy varied by statement type.

  • Licensing matters commercially: AlphaFold 3 weights and the AlphaGenome API carry non-commercial terms.

  • Provenance, metadata and negative results are the scarce inputs for closed-loop discovery.


Actionable Next Steps


  1. Read the current criteria at millenniumproblems.bio and note the version date before citing any problem.

  2. Scientists: map your work against the exact preregistration requirements and state which half of a problem you address.

  3. Biotech leaders: audit licences for AlphaFold 3, AlphaGenome and any agent platform before building on them.

  4. AI-for-science teams: report prospective, blinded hit rates and log which steps the AI performed.

  5. Lab leaders: standardise assay formats and capture ELN and LIMS metadata, including negative results.

  6. Investors: ask for independent replication and cost per validated hit, and treat preprints as preprints.

  7. Follow official updates on adjudication and prizes, and use the site’s comment form to suggest corrections.

  8. Complete biosafety and ethics review before any closed-loop work involving vectors, cell-entry proteins or animals.


Glossary


  • AI scientist: a software system that generates hypotheses, analyses data and plans experiments.

  • Active learning: choosing the next experiments to maximise information gained.

  • Protein language model: a model trained on protein sequences to predict or generate them.

  • De novo protein design: creating proteins not derived from natural sequences.

  • Zero-shot: working on a new target without target-specific training or optimisation.

  • Rubisco: the enzyme that fixes CO₂ in photosynthesis.

  • kcat: an enzyme’s turnover number, or reactions per second per active site.

  • Specificity (Sc/o): Rubisco’s preference for CO₂ over O₂.

  • Codon: a group of bases specifying an amino acid.

  • Quadruplet codon: a four-base codon.

  • Protease: an enzyme that cuts proteins.

  • Protein binder: a protein designed to attach to a target molecule.

  • AAV: adeno-associated virus, a common gene-therapy vector.

  • Lentivirus: an enveloped virus used as a gene-therapy vector.

  • Vitrification: cooling into a glass-like state without ice.

  • Nanowarming: rewarming by heating nanoparticles with magnetic fields.

  • Polymerase: an enzyme that builds nucleic acid chains.

  • Reverse transcriptase: an enzyme that copies RNA into DNA.

  • RdRP: RNA-dependent RNA polymerase, which copies RNA from RNA.

  • Nitrogenase: the enzyme that converts N₂ to ammonia.

  • Wet lab: a laboratory doing physical experiments.

  • DBTL: design-build-test-learn, an iterative engineering loop.


Sources & References


 
 
bottom of page