52.1% vs 35.6%: what the pharma AI federation proves — and what it doesn't

Five competing drug companies — AbbVie, Astex Pharmaceuticals, Bristol Myers Squibb, Johnson & Johnson and Takeda — together with a research group at Columbia University have fine-tuned a shared AI model on more than 20,000 proprietary…

Illustration: five sealed glass vials holding different colored liquids are connected by thin metal tubes merging into one central laboratory instrument, a metaphor for federated AI learning where competing drug companies share one model without sharing their raw data.
Illustration
Gift article

52.1% vs 35.6%: what the pharma AI federation proves — and what it doesn't

Five competing drug companies — AbbVie, Astex Pharmaceuticals, Bristol Myers Squibb, Johnson & Johnson and Takeda — together with a research group at Columbia University have fine-tuned a shared AI model on more than 20,000 proprietary molecular structures without a single structure file leaving the companies' own systems. The model, called AISB-1-Fed, reportedly beat both the public base model and the strongest open reference model on a held-out test set of private structures, according to the consortium itself. The entire run took under ten weeks, and the result was announced in the week around September 23, 2026 — at the same time as the same network launched a new initiative on binding affinity.

All the figures, however, come from the consortium itself, relayed through secondary coverage, and the result has not been peer-reviewed. It is, in other words, a remarkable proof of concept — with clear caveats.

The reported result

The concrete claim is this: On a held-out benchmark of 1,056 private structures, AISB-1-Fed achieved high-quality predictions of the interfaces between protein and ligand molecule for 52.1 percent of the structures. By comparison, the public base model, OpenFold3 Preview 2, reached 35.6 percent, while Boltz-2 — which according to the coverage was previously considered the strongest public reference model on this type of task — scored 40.9 percent (TechTimes; MSN).

The jump from 35.6 to 52.1 percent is not marginal: it means the model, according to the consortium, solves roughly 16.5 percentage points more of the test tasks. But it is also worth noting that the test set consists of private structures — precisely the kind of data the model was trained on via the participating companies, set aside as held-out. That makes the comparison with public models partially uneven from the outset: the competing models have never seen this type of pharmaceutical complex, while AISB-1-Fed has been trained on data from the same domains. The model is better adapted to the task type — which is the point of the project, but also part of the explanation for the numbers.

How federated learning worked here

The mechanism is federated learning, a technique in which the data stays where it is and the model travels instead. In practice:

  1. Each company received a copy of the starting model, OpenFold3 Preview 2, running inside its own security environment.
  2. The copy was further trained on the company's own, proprietary protein–ligand structures, behind the company's own firewalls.
  3. Only the model updates — the weight changes representing what the model learned — were sent back and combined into a shared model. No structure file left any company along the way, according to MSN.

The entire run took place across five corporate environments on three continents and was completed in under ten weeks. The orchestration of the federation itself was powered by the company Apheris (TechTimes). Apheris CEO Robin Röhm is cited in the coverage.

Technically, important questions remain that the coverage does not answer in detail. Federated learning protects raw data, but the model updates themselves can in principle leak information about the training data. Whether the consortium used additional protective techniques on the updates is not clear from the public account. This is an open question that should be answered before the model can be presented as full proof of privacy preservation.

Who participated, and when

The collaboration takes place under the umbrella of the AISB Network. AbbVie and Johnson & Johnson were founding members; Astex Pharmaceuticals, Bristol Myers Squibb and Takeda joined in October 2025. According to TechTimes' account, this means the initiative includes five of the 20 largest pharmaceutical companies globally (TechTimes).

The contributions were not evenly distributed. AbbVie alone contributed more than 9,000 protein–ligand structures — a substantial share of the combined dataset. In total, the companies contributed structural data from more than 20,000 experimentally determined protein–small molecule complexes drawn from active drug development programs. According to the consortium's own figures, this roughly triples the amount of drug-relevant structural data available for model training compared with what exists in public repositories (TechTimes).

The research group of Mohammed AlQuraishi at Columbia University participated as a scientific partner, with OpenFold3 Preview 2 as the starting model on which AISB-1-Fed was built, according to the coverage. The companies contributed the data, Apheris the infrastructure. All parties have an interest in the result appearing to be a success, and it is worth keeping that in mind when weighing the numbers.

Why this is more than a technical demonstration

What is formally new here is not federated learning as a technique. What is new is that five direct competitors, each holding structures produced through extremely costly experiments, found it worthwhile to do this together. That is a regime question, not a technical one: pharmaceutical structural data has traditionally been the most guarded data category in the industry, because it is the product of decades of medicinal chemistry and is expensive to produce.

The value proposition is asymmetric in an interesting way: each company gives up exclusive access to its own data advantage in exchange for a shared model that is better than anything it could have built alone. The tripling of training data is the mechanism behind the 52.1 percent figure — none of the participants could have reached it alone, and none of them gives away their raw data. That is what makes the result a playable template for other industries where the best data is locked away with competitors: banks with fraud data, hospitals with patient records, insurers with claims files.

That the collaborative model still has life in it suggests the participants themselves see more value in the model: the same week the result was announced, the AISB Network launched a follow-on initiative aimed at the next big problem in early drug development — predicting precisely how tightly a molecule grips its target, in other words binding affinity (TechTimes).

The caveats — and what it would take for this to count as more than self-evaluation

It is important to be honest about what exists here. First: the benchmark is internal. The figures 52.1, 35.6 and 40.9 percent originate from the consortium itself, are not peer-reviewed, and the test set consists of the consortium's own held-out structures. No independent party has repeated the measurement. Before this can be counted as a scientific result rather than a successfully self-reported project, one would want a published methodology with details about the test set, how "high quality" of the interface predictions is defined and measured, and how the public baseline models were configured and run — along with the possibility of independent replication.

Second: a better interface prediction number is not a drug. No approved — or even candidate — drugs have been delivered as a result of AISB-1-Fed, and there is no documentation that better predictions of this type shorten real drug development timelines. The distance between a higher benchmark percentage and an approved drug is measured in years and significant sums.

Third: the privacy mechanism is thinly documented in public coverage. The claim that no structure file left any company is credible as a description of the data flow in federated learning, but the technical details of how the model updates were protected against indirect leakage remain an open question. For competitors considering copying the model, that is precisely the question they must ask.

Finally: the value distribution is unclear. AbbVie contributed more than 9,000 structures — what did AbbVie get in return for competitors gaining access to a model trained on AbbVie's data? Such distribution and ownership questions must have been negotiated, but the details are not public.

The most cautious — and most plausible — view is this: even if the numbers hold up under independent scrutiny, the main result is not the model but the proof that the collaborative model works in practice. Five competitors managed to coordinate data contribution, model training and security architecture in under ten weeks without anyone giving away their crown-jewel material. If that pattern spreads to finance, healthcare and other tightly guarded data sectors, it will be far more consequential than one benchmark number from one consortium.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.