In a transformative move for the future of medical diagnostics, Cognita Imaging—a startup rapidly rising to prominence in the health-tech sector—has secured a $1.29 million research contract from the U.S. Food and Drug Administration (FDA). This partnership marks a critical milestone in the government’s attempt to reconcile the rapid, open-ended evolution of generative artificial intelligence (AI) with the rigorous safety standards required for clinical healthcare.
The project aims to pioneer a novel evaluation framework: utilizing a "jury" of large language models (LLMs) to audit the diagnostic accuracy of AI-generated radiology reports. By automating the oversight process, the initiative seeks to solve one of the most persistent bottlenecks in modern medicine—the time-consuming, labor-intensive validation of AI tools that are becoming increasingly complex.
The Problem: Moving Beyond Single-Task AI
For years, the FDA’s approach to AI in radiology has focused on "narrow" AI—specialized algorithms designed for a single, binary task. These tools are often highly efficient at detecting specific anomalies, such as a pulmonary embolism in a CT scan or a fracture in a skeletal X-ray.
"Evaluating these types of models is straightforward because they are designed for specific, constrained tasks," explains Akshay Chaudhari, co-founder of Cognita and an associate professor of radiology and biomedical data science at Stanford University. "However, tools that only identify one or two things are not particularly helpful in the day-to-day workflow of a radiologist."
The clinical reality is that a single scan often contains hundreds of potential findings. Newer, more sophisticated generative AI models are being engineered to interpret these images and generate comprehensive, full-text clinical reports. This transition from "one-off" detection to holistic report generation has created a regulatory vacuum. When an AI can describe hundreds of findings in an open-ended format, the traditional method of having human experts manually cross-check every word becomes mathematically impossible to scale.
"It’s more open-ended in nature," Chaudhari notes. "So, how do you meaningfully evaluate these hundreds of findings at scale?"
Chronology: From Foundation to Federal Partnership
The trajectory of Cognita Imaging has been marked by rapid acceleration and strategic integration within the U.S. healthcare system:
- 2024: Cognita Imaging is founded, with a mission to bridge the gap between advanced deep learning and clinical radiology workflows.
- 2025: Only one year after its inception, the startup is acquired by Radiology Partners, the largest physician-led radiology practice in the United States. This acquisition provides Cognita with an unprecedented feedback loop, allowing its developers to test their models against real-world data from diverse settings, ranging from academic medical centers to rural outpatient clinics.
- March 2026: The FDA grants "Breakthrough Device Designation" to Cognita’s vision-language model, which is designed to interpret chest X-rays and draft preliminary reports.
- June 22, 2026: Cognita commences an 18-month research contract with the FDA’s Center for Devices and Radiological Health (CDRH).
- September 2026: The $1.29 million contract is formally announced, signaling a broader federal interest in codifying the regulatory framework for generative AI.
The "LLM Jury" Methodology
The core of the new FDA-funded project is the development of an automated oversight system. Cognita is building a framework where a panel of disparate LLMs acts as a "jury" to evaluate the output of another AI model.
The methodology is designed to be rigorous and data-intensive. Cognita will apply this evaluation framework to a massive dataset of 1 million patient exams, sourced from a diverse U.S. cohort. The research will specifically measure performance variations across:
- Different patient demographics and health backgrounds.
- Various clinical care settings (e.g., emergency departments vs. elective outpatient clinics).
- Multiple types of imaging hardware and manufacturers.
- Rare disease states that often baffle less-sophisticated AI models.
To ensure the "jury" itself is reliable, the project includes a failsafe mechanism: when the LLM jury encounters a disagreement of clinical significance, the system triggers an alert for human review. Expert radiologists will then investigate whether the discrepancy originated from the initial AI’s analysis, the LLM jury’s interpretation, or the original, human-written baseline report.
Addressing the "Hallucination" Crisis
One of the primary anxieties surrounding generative AI in clinical settings is the phenomenon of "hallucinations"—instances where an AI generates content that sounds medically plausible but is factually incorrect or entirely fabricated. Equally concerning are "omissions," where a model fails to report a critical finding, such as a subtle tumor or an early-stage infection.
Cognita’s project aims to provide the first quantifiable metrics for these risks. By running the jury system against a large-scale cohort, the team intends to establish a baseline for how often these errors occur and, crucially, whether they are "clinically meaningful." The goal is to move beyond abstract worries about AI error and into a realm of statistical certainty where regulators can set tolerance thresholds for clinical deployment.
Official Responses and Regulatory Implications
The FDA has been increasingly active in its oversight of medical AI. Last month, the agency released a discussion paper outlining its current thinking on generative AI in medical devices, signaling that it is actively seeking to evolve its regulatory posture.
Chaudhari emphasizes that the ultimate goal of the Cognita project is not merely to get a specific product approved, but to establish a "generalizable" framework that the FDA can integrate into its long-term regulatory guidance. "We’re confident that through these very nice partnerships that the FDA fosters, there is a path forward," he said.
The deliverables for the FDA at the conclusion of the 18-month contract are comprehensive:
- Software Code: The underlying architecture for the LLM jury framework.
- Guidance Documents: Best practices for developers on how to build and validate LLM-based evaluation juries.
- Comparative Analysis: Data comparing the efficacy of the framework across large versus small validation cohorts.
- Discrepancy Studies: A detailed log of human-radiologist-reviewed errors and disagreements.
The Clinical Necessity: Combatting the Radiologist Shortage
The urgency of this project is underscored by a deepening crisis in the American healthcare workforce. The United States is currently grappling with a significant shortfall of radiologists. As the population ages and the demand for diagnostic imaging rises, the remaining workforce is facing unsustainable workloads, longer hours, and massive case backlogs.
These backlogs create a "delay-of-care" risk that can be as dangerous as a misdiagnosis. By automating the generation of preliminary reports, tools like those being developed by Cognita could act as a force multiplier for radiologists, allowing them to spend their time verifying and refining AI findings rather than typing out repetitive, standard-case reports from scratch.
Looking Toward the Future: The Breakthrough Designation
While the $1.29 million research contract focuses on the evaluation of AI, Cognita’s "Breakthrough Device Designation" remains a parallel, yet distinct, effort to bring their specific vision-language model to market.
This designation—reserved for medical devices that provide for more effective treatment or diagnosis of life-threatening or irreversibly debilitating human diseases—is a testament to the potential the FDA sees in Cognita’s work. The company is currently engaged in ongoing, high-level discussions with the agency to design the clinical studies required for final authorization.
"There’s still work to be done in developing the exact study plan and actually implementing the study," Chaudhari noted. "But we recognize that there is a real, pressing clinical need for these tools."
As the integration of generative AI into radiology transitions from theoretical research to clinical reality, the work being conducted by Cognita Imaging represents a critical bridge. By teaching machines to audit machines, the healthcare industry may finally possess the tools to scale high-quality diagnostic care, even as the shortage of human radiologists threatens the stability of the system. The "jury of LLMs" may not replace the doctor, but it may very well be the key to ensuring that the doctor has the time and the accurate data needed to save lives.
