FDA funds LLM jury to grade AI radiology reports across one million exams
An 18-month research contract will test whether multiple language models grading radiology reports in parallel can measure the quality of autonomous AI systems better than human reader studies of a few hundred cases.
The U.S. Food and Drug Administration has awarded the company Cognita Imaging a research contract worth $1.29 million to build and validate a so-called "LLMs-as-a-jury" framework: a system in which multiple large language models simultaneously evaluate radiology reports written both by humans and by generative AI. The contract runs for 18 months, took effect on 22 June 2026, and is funded under the FDA's Broad Agency Announcement program for regulatory science research and development. Trade publication HIT Consultant reported the news on 16 September 2026 and is so far the only source for the contract details.
The project is titled "Virtual Subject Matter Expert Agents for Radiology Report Evaluation" and is led by Dr. Akshay Chaudhari, co-founder of Cognita and associate professor of radiology and biomedical data science at Stanford University. Co-founder and CEO Dr. Louis Blankemeier serves as co-investigator. Cognita is a wholly owned subsidiary of Mosaic Clinical Technologies and operates within the broader Radiology Partners ecosystem. The contract is a research assignment, not an approval: according to HIT Consultant's account, it is meant to "advance regulatory science standards" and "does not constitute formal FDA product approval or commercial clearance."
The problem the contract targets
Medical AI approval has traditionally rested on reader studies, in which radiologists assess a tool's output against a ground truth. Such studies typically include only a few hundred cases. That means rare edge cases, subtle discrepancies between equipment vendors, and regional variations in workflow barely register — precisely as autonomous generative AI systems that write radiology reports without human pre-review approach deployment in hospital settings. The project's rationale, as HIT Consultant describes it, is exactly this gap between small evaluations and large, varied patient populations.
How the LLM jury works
The framework Cognita is to build has several large language models working simultaneously on the same task: examining and cross-examining both human-written and AI-generated imaging reports. The point of using multiple models is to reduce bias, so that a single model's blind spots do not alone determine the verdict — the same logic as a jury, where multiple votes are meant to counteract one person's bias.
The system builds on GREEN (Generative Radiologist Evaluation and Error Notation), an open-source framework Cognita developed earlier to capture clinically meaningful deviations between the radiologist's ground-truth report and AI-generated text. The jury thus does not stand free: it is measured against reports written by human radiologists.
A key feature is human escalation. The platform isolates high-discrepancy outliers and clinically significant disagreements and routes them to expert review. Radiologists then determine where the error originates: in the generation model that wrote the report, in disagreement within the LLM jury itself, or in ambiguity in the original human-written ground truth. This three-way split is essential for regulatory use, because a low score does not necessarily mean the AI system failed — the jury or the reference standard may be at fault instead.
The scale and validation plan
The evaluation system is, according to the plan, to be applied to roughly one million patient exams from a diverse U.S. cohort. The analysis is intended to map performance variation across demographics, care settings, imaging equipment manufacturers, and low-prevalence pathologies. In addition, the contract includes a sample-size sensitivity analysis to quantify how many edge-case failures small evaluation datasets actually miss — figures that could justify why reader studies alone are not sufficient.
The initiative builds on clinical workflows and data pipelines from MosaicOS and the Radiology Partners network, which the companies themselves say cover more than 4,000 radiologists and interpret more than 55 million exams annually. Those scale figures are the ecosystem's own claims, not independently verified. The details of the division of responsibilities among Cognita, Mosaic Clinical Technologies and Radiology Partners are not described in the available source material.
Deliverables — and their limits
Cognita is to deliver three things to the FDA: open-source software code, guidelines for benchmarking multi-agent LLM juries, and comparative performance analyses usable both in premarket review and postmarket surveillance.
The limits matter as much as the deliverables. This is regulatory science — methods development for the regulator — not approval of any product. And more fundamentally: everything that exists so far is plans. The method has not been validated for regulatory decisions, and the effectiveness of the LLM-jury approach is unresolved. One million exams is a target, not a fact; neither data access, privacy arrangements nor ethical approval for the dataset are documented in the source material.
Open questions
Three issues bear watching. First, all contract facts — the amount, dates and dataset size — currently rest on a single secondary source, HIT Consultant. Neither the FDA's own announcement, public contract registries nor a Cognita press release are available in the material, so the figures cannot be cross-checked.
Second, the data flow is unsettled: how would roughly one million patient exams from a commercial radiology network flow into an FDA-funded research project, and with what privacy guarantees? That question matters far beyond this project, since similar arrangements could become the template for future medical AI regulation.
Third, it remains unanswered whether an LLM jury is actually reliable enough to enter regulatory decisions. Disagreements among jury members must be interpreted by humans — something the project itself anticipates through its escalation mechanism. Whether the jury can in practice distinguish an AI system's errors from ground-truth ambiguity is exactly what the 18-month contract is meant to find out.

