Federated training beat a centralized baseline on chest X-rays, by about two percentage points of Macro-AUC
A study published in the journal Machine Learning, reported by Bioengineer.org on 25 September 2026, found that federated training not only protected patient images but also outscored centralized training on chest X-ray classification. The difference is modest – about two percentage points in Macro-AUC – and all the figures come from simulated experiments on a public dataset, not from real hospitals. But if the finding holds, it challenges a long-standing assumption in medical AI: that data must be gathered in one place to get the best model.
What is new
The research group behind the study is led by Suresh Arumugam of Dayananda Sagar University in Bengaluru, with colleagues from three other Indian institutions. They introduce FedXAI-Health, a federated learning framework aimed at multi-label classification of thoracic diseases on X-rays – that is, detecting multiple possible findings in the same image (Bioengineer.org).
What is new in the reporting is not the idea of federated learning itself, which has long been proposed as an answer to privacy requirements in healthcare. It is the claimed result: that under severe non-uniform data conditions – where each "hospital client" has a skewed and different distribution of images – the federated model achieved better scores than a centralized model trained on all the data pooled together. That advantage is counterintuitive, and so far it is documented only through secondary reporting of the authors' own results.
How federated learning works
The principle is to send the model to the data, not the data to the model. Each participating hospital client trains a shared neural network on its own machines, locally and on its own images. Only the model updates – not the patient images – are then sent to a central server, which merges them into an updated common model. The images never leave the client (Bioengineer.org).
For hospitals subject to privacy regulation such as HIPAA in the United States and GDPR in Europe, this is more than a technical detail. According to the reporting, such laws make it nearly impossible for medical institutions to pool their X-rays into a single training set, meaning each institution's model learns only from its own, often narrow, patient population. A setup in which images stay local can therefore enable collaboration on medical models where data pooling would be blocked.
The algorithm used to aggregate the updates is FedAvg, the classic federated averaging algorithm. The network's backbone is EfficientNetB0, according to the reporting. The framework is thus built on well-known, established components.
The numbers – and how they were produced
The experiment is a simulation, not a deployment. The team created a network of five simulated hospital clients, each client receiving a differently distributed share of the publicly available NIH Chest X-ray 14 dataset, with multi-label classification of thoracic findings as the task (Bioengineer.org).
To mimic the real-world problem – that different hospitals see different patient populations – the data were partitioned using Dirichlet distributions with concentration parameters from 0.1 to 5.0. A parameter of 0.1 yields extremely skewed client datasets; a parameter of 5.0 approaches an even distribution. According to the reporting, the team validated robustness across this entire range.
The headline result, as reported: FedAvg with EfficientNetB0 achieved a Macro-AUC of 0.8060 ± 0.0048, while the centralized baseline model – trained on all the data pooled together – managed 0.7859 ± 0.0142. The source states the difference as "2.0 percent"; the numbers indicate this amounts to roughly two percentage points (0.8060 − 0.7859 = 0.0201) in favor of the federated approach (Bioengineer.org).
Macro-AUC measures how well the model distinguishes positive from negative cases across all the diagnostic labels, on average. All of these figures rest on the secondary source's rendering of the authors' results.
Why the result is surprising
The common expectation in the field has long been that federated learning is a trade-off: you get privacy, but pay with lower model performance, because skewed local data produce skewed local updates that are harder to weigh than a pooled training process.
That FedAvg in this study beat the baseline outright is therefore not just a practical detail – it suggests that something in the setup confers a real advantage, rather than merely avoiding a loss. The reporting, however, does not convey any complete explanation from the authors of why the federated model won. The counterintuitive advantage thus stands as a claim requiring verification against the primary article, not as an explained finding.
Key caveats
The most important limitation is that this took place in simulation on a public dataset, with five constructed clients – not in a live hospital network. The artificial Dirichlet partitions capture only certain forms of skew. Real hospital differences also include different scanners, protocols, labeling practices, and patient groups, and how the framework would fare in that reality was not tested here.
In addition, all the figures rest on Bioengineer.org's coverage of the study in Machine Learning; the findings should therefore be read as the authors' own results as reported, not as independently verified results.
What it could mean
If the results hold up under replication, they have implications beyond a single experiment. They suggest that privacy-preserving collaborative training is not necessarily a compromise solution – and that the assumption that data pooling is necessary for the best possible model may be worth retesting. For regulated hospital environments under HIPAA and GDPR, lowering the barrier to multi-institutional model training could mean access to more varied training data, and thus models that generalize better across patient populations.
Open questions
Three things are missing before this can be considered settled. First: an explanation of why federated training beat centralized training, which the reporting does not fully convey. Second: independent replication of the figures, ideally by groups unaffiliated with the study. Third: performance in real hospital networks, with real clients, real networks, and real regulatory requirements – conditions the simulation only hints at. Until then, the most reasonable position is that the team has reported an interesting result which, if confirmed, will be worth taking seriously.

