A coalition led by Biohub, the US Department of Energy and the National Institutes of Health announced a $1.8 billion commitment to generate and openly share the data needed to train predictive AI models of biology — what researchers call the pursuit of a "virtual cell."

The announcement, made Wednesday in Redwood City, California, brings together federal agencies, philanthropy and three of the world's most prominent AI organizations in what Biohub describes as the largest coordinated commitment to generating AI-ready biological data to date. For more context on this story, see our ongoing latest AI developments.

Who Is Paying For What

The $1.8 billion figure combines money, data, computation and new measurement technology from several partners:

  • Biohub anchors the effort with its founding $500 million commitment. Of that, $400 million supports new measurement technologies that expand what biologists can actually observe: cryo-electron tomography, which resolves near-atomic detail inside the cell, alongside advanced microscopy designed to image cells at scales and speeds that current laboratory workflows cannot approach.
  • The Department of Energy will invest more than $500 million over five years in laboratory measurement, modeling and computation through its Office of Science.
  • The NIH will coordinate the contribution of datasets, repositories and knowledge bases developed through more than $500 million in prior federal investment, with Biohub working alongside the agency to standardize those datasets for AI model training.
  • Google DeepMind, Isomorphic Labs and Meta are collectively investing $300 million in the Virtual Biology Initiative, the international effort — first announced in April 2026 — that today's expansion formalizes.

The result will be an open resource for the global research community, not a proprietary dataset locked inside any single company.

An Expansion, Not a Beginning

Today's announcement formalizes and scales up the Virtual Biology Initiative, which was first announced in April 2026 with a mandate to coordinate data generation across institutions and scientific disciplines. The April launch established the framework; the new commitments fill it with federal money, federal data and private-sector investment.

The division of labor matters. The DOE brings world-class computing capacity and national laboratory instrumentation. The NIH brings standardized datasets accumulated through decades of federally funded research — the kind of curated, longitudinal biological data that no single lab or company could regenerate. Biohub, the research organization backed by tech philanthropy, acts as the integrating hub that will normalize these sources into formats AI models can actually train on.

The 'Virtual Cell' Grand Challenge

The initiative's stated goal is to make it possible for scientists to run experiments digitally: to predict how a cell responds to a drug, a genetic perturbation or a disease state before touching a pipette.

"An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally," said Alex Rives, Biohub's Head of Science, in the announcement. "The insights that come from this could unlock a far greater understanding of disease and open up completely new paths for cures."

"Because of this potential, the creation of a virtual cell is one of the most important challenges for the next era of science," Rives added. "It will require coordinated data generation efforts at a national and international scale, which is why these partners are coming together."

Why Data, Not Just Models, Is the Bottleneck

AI in biology has produced landmark results — DeepMind's AlphaFold family demonstrated that neural networks can predict protein structures with experimental accuracy — but those systems were trained on decades of curated public data. The frontier problems, from predicting how entire cells respond to interventions to modeling cellular interactions, require data that largely does not yet exist in AI-ready form.

That is the gap the initiative targets. Rather than releasing another model, the partners are funding the unglamorous infrastructure of science: expanding cell response data to interventions across far more cell types and conditions than have yet been studied, building and validating technologies for studying cells and their interactions at greater scale, speed and accuracy, and assembling shared repositories that outside labs can actually use.

The expansion also has a defensive dimension. On Hacker News, where the announcement drew significant attention, commenters contrasted the open-data commitment with the current US administration's removal of previously public research datasets from government servers — a tension between open science at the frontier and retrenchment elsewhere.

What It Means for AI in Drug Discovery

For the pharmaceutical and biotech industry, the message is that foundational data is becoming a shared public utility. Isomorphic Labs, Alphabet's drug design company, and Meta — which has its own fundamental AI research and protein-structure work — are both betting that better shared data will accelerate everyone's models, including their own.

If the initiative succeeds, the near-term deliverables will be expanded cell-response datasets covering vastly more cell types and conditions, plus validated measurement technology that shortens the path from hypothesis to high-quality training data. The long-term prize is a simulation good enough to triage drug candidates digitally before the first wet-lab experiment.

The scientific community is being explicitly invited to participate: Biohub's announcement closes with an open call for the "worldwide scientific community" to join the project.

For more on how machine learning is reshaping the life sciences, see AI Buzz Wire's AI research coverage.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →