Biohub coalition commits $1.8 billion to build the data layer for virtual cells
Biohub, US science agencies, Google and Meta are combining measurement technology, computing and biological datasets to train predictive cell models—but a reliable virtual cell remains a target, not a finished product.

The story
A coalition spanning Biohub, two US science agencies and major AI companies has assembled what it describes as a $1.8 billion commitment to build the experimental data needed for predictive models of human cells. Announced on October 7, the expanded Virtual Biology Initiative brings together the US Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs, Meta, research institutes and technology suppliers. The immediate product is not a drug or a working digital cell. It is a coordinated measurement and data infrastructure intended to make those advances more plausible.
The headline number needs careful parsing. Biohub committed $500 million over five years when it launched the initiative in April. DOE now plans to contribute more than $500 million over five years for biological measurement, modeling and computation. Google DeepMind, Isomorphic Labs and Meta are jointly committing $300 million. NIH will coordinate datasets, repositories and programs created through more than $500 million in earlier federal investment. The $1.8 billion therefore combines new funding with existing data, computing and scientific infrastructure; it is not simply a new cash pool deposited in one account.
The scientific problem is scale and consistency. Modern experiments can measure which genes are active in a cell, where molecules sit inside tissue and how cells change after a genetic or environmental perturbation. Yet most datasets were built for individual questions, on different instruments and under different protocols. Language models benefited from vast stores of digitized text. Biology has no comparable, standardized corpus that captures cells across enough conditions to teach a model the rules of their behavior.
The initiative proposes to generate that missing layer through spatial transcriptomics, high-throughput and high-content microscopy, cryo-electron tomography, whole-body atlases and systematic perturbation screens. DOE's national laboratories add exascale computing, advanced imaging and autonomous laboratory capabilities. NIH brings biomedical repositories and its Bio Genesis Mission. Participating organizations also include the Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute, while NVIDIA is listed as a supporting technology partner.
The ambition is a virtual cell: a model able to predict how a cell will respond when a gene is switched off, a drug is added or its environment changes. That is a higher bar than classifying an image or finding a correlation in an existing dataset. A useful model must make prospective predictions that survive experiments on cells it has not already seen. Biohub's head of science, Alex Rives, told Reuters that today's cell datasets contain hundreds of millions of cells, while accurate predictive models may require billions and eventually trillions of observations. The partners aim to produce a first dataset in about a year and useful predictive models within five years, but those are goals rather than demonstrated results.
Access will be a defining test of the program's open-science claim. Biohub says the coalition will create an open resource with common standards and broad access. Reuters reported that commercial funders will receive an embargo period on datasets they finance before those datasets are released publicly. Rives said the parallel government-funded work will carry no such restriction. Temporary exclusivity may attract private capital, but the value of a shared biological foundation will ultimately depend on how quickly outside researchers can inspect, reproduce and challenge its data.
INNOVOX analysis: this initiative moves the center of gravity in AI biology from model architecture to measurement infrastructure. Compute can be purchased and algorithms can be copied, but carefully controlled experiments remain slow, expensive and dependent on specialized instruments. A $1.8 billion coalition can generate far more observations than a typical academic laboratory, yet volume alone will not solve biology. The critical asset will be metadata: precise records of sample origin, experimental conditions, instrument settings, perturbations and uncertainty. Without that discipline, a model may learn laboratory-specific artifacts instead of transferable biology.
The initiative could also change how discoveries are validated. If a model predicts that a specific perturbation will reverse a disease-related cellular state, autonomous laboratories could test the claim, feed the result back into the model and select the next experiment. That closed loop could reduce wasted screening and reveal mechanisms that conventional trial-and-error misses. It does not eliminate animal studies, clinical trials or regulatory evidence, and it cannot guarantee shorter drug-development timelines. The proof will come from prospective benchmarks: predictions made in advance, tested across laboratories and compared with strong biological baselines.
What to watch now is not another funding total but the first release. Its license, embargo, documentation and representation of diverse donors and cell states will show whether the coalition is building shared infrastructure or a collection of privileged data silos. Researchers should also look for independent replication and for models that generalize across instruments and institutions. A credible virtual cell will be measured by the experiments it gets right—not by how convincingly it describes biology after the fact.
INNOVOX analysis
The coalition is treating experimental data—not a larger neural network—as the scarce infrastructure for AI biology. Its success will depend on whether measurements made across different laboratories, instruments and cell systems can be standardized well enough for models to predict unseen interventions rather than merely reconstruct familiar patterns.
What to watch
Watch for the first dataset's release terms and benchmark results, especially predictions tested prospectively in wet laboratories. Also track common metadata standards, cross-lab reproducibility, privacy and consent rules for human-derived samples, and the duration of commercial embargoes before funded data becomes public.
