Biohub to standardize NIH data for AI training in new $1.8 billion push

On Wednesday, October 7, 2026, the nonprofit Biohub announced that the US Department of Energy, NIH, Meta, Google DeepMind and Isomorphic Labs are joining the organization's Virtual Biology Initiative – lifting total commitments to $1.8…

Illustration: laboratory samples of varying shapes and colors being sorted into a precise grid of identical slots, a metaphor for standardizing biological data.
Illustration
Gift article

Biohub to standardize NIH data for AI training in new $1.8 billion push

On Wednesday, October 7, 2026, the nonprofit Biohub announced that the US Department of Energy, NIH, Meta, Google DeepMind and Isomorphic Labs are joining the organization's Virtual Biology Initiative – lifting total commitments to $1.8 billion, according to Biohub's own figures given to Reuters. The goal is to build open training datasets for AI models in biology.

But the announcement carries a tension worth examining closely: the initiative is marketed as open science, with the datasets to be a shared resource for the entire research community. At the same time, commercial funders receive an embargo period – about one year, according to Biohub's chief scientist Alex Rives – during which they can work with the data they funded before it is made publicly available. In other words, the same data Biohub says is necessary to train a "universal virtual cell" first passes through a commercial gate.

Who is contributing what

The $1.8 billion figure is Biohub's own stated sum, reported by Reuters (Krystal Hu), and it consists of very different elements:

  • Meta, Google DeepMind and drug discovery company Isomorphic Labs are investing a combined $300 million.
  • The Department of Energy (DOE) is investing over $500 million spread across five years, in laboratory measurements, modeling and computation.
  • NIH is coordinating datasets and databases built with over $500 million in prior federal funding, which Biohub will standardize for AI training.
  • Biohub itself contributed $500 million when the initiative launched in April.

It is worth emphasizing that this is not $1.8 billion in new cash. The NIH portion consists of resources that are already funded, now being coordinated and standardized – valuable, but of a different character than new investment. The available sources also do not explain how the pieces sum precisely to the $1.8 billion total, and no primary documents from Biohub, DOE or NIH are available in the material at hand. All figures should therefore be read as Biohub's own statements, relayed through Reuters and Axios.

How the data arrangement works

The core idea is simple: Biohub collects and standardizes biological cell data at large scale and makes it available for AI training. What's interesting is the timing of the distribution.

"With commercial funders, we have embargo periods, where there is a period during which the groups can work with the data, before it becomes available as a public scientific resource," Biohub's chief scientist Alex Rives told Reuters.

To Axios he was more specific, describing the length this way: "We have to have some incentive for commercial actors to participate in this, and the embargo period creates that. But it's a one-year embargo. So that means the data quickly becomes broadly available for scientific efforts."

Publicly funded research will, according to Rives, not be subject to such restrictions. The logic is that a one-year head start – time to work with fresh data before anyone else – should be incentive enough for companies like Meta, Alphabet-owned DeepMind and Isomorphic Labs to put up capital. But no policy or contract documentation describing the terms in detail is yet available; what we know about the embargo comes from Rives's own descriptions to journalists.

The scientific goal – and how hypothetical it is

According to Rives, today's datasets consist of hundreds of millions of cells, while an accurate predictive model requires billions and eventually trillions. Closing that gap is the whole point of the initiative.

The first phase is to build a broad map of cellular biology, gathering different types of information about cells and how they respond to changes. Further out, according to Rives, an accurate predictive model for biology could dramatically accelerate scientific discovery: "An accurate predictive model for biology could dramatically accelerate scientific discovery by enabling researchers to perform experiments digitally," he said, according to Dow Jones Newswires. Axios describes his vision as models that can predict the molecular causes of individual diseases – a kind of "digital experiments" replacing some laboratory work.

Here a sharp distinction must be drawn between claims and projections. The timelines Rives lays out – a first dataset within roughly a year, predictive models within five years – are his own projections, not documented milestones. And the fundamental scientific question remains unresolved: whether biology follows useful scaling laws at all, the way language models have with text. Axios itself flagged this as an open question. If cell biology does not scale predictably with data volume, billions and trillions of cells may end up buying less than hoped.

Open science, with a footnote

Priscilla Chan, Biohub's co-founder, defended the open profile in a Reuters interview: "Biology has so far been a kind of clever, discovery-based science," she said. "We've always treated this as a community resource, not just for one group, so that it can build on itself over time."

It is the contradiction that makes the story worth reading: the datasets are described as a community asset that will build value over time – yet initially they are, for the funding companies, a competitive advantage for up to twelve months. One year is a short time at AI development pace, and Rives's point that the data quickly becomes "broadly available" is real. But the ordering matters: the group that funds a dataset gets to work with it before anyone else. Whether that is a reasonable price for bringing private capital into public research infrastructure, or a dilution of the open-science ideal, is a value judgment – not something the sources answer.

The context: the race for biology data

The announcement lands in the middle of a clear pattern. The major AI labs are buying into biological data generation: Anthropic has pressed further into biology with a wet laboratory, while the OpenAI Foundation has launched a grant program of over $125 million to fund biological and medical datasets for AI research. The Biohub deal stands out by involving public institutions so directly, but common to all these initiatives is the recognition that the next generation of bio-AI models needs data that does not yet exist.

What is worth watching

The most obvious things to follow are two. First: the first dataset is expected within about a year, according to Rives – a timeline that can be checked relatively quickly. Second: whether the embargo arrangement holds up in practice as a genuine open-science model, or whether a one-year commercial head start becomes a precedent that other major research initiatives copy – and whether the terms are ever documented publicly, so they can be held to account.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.