An AI assistant for retinal disease diagnosis posts strong trial results
A multicenter randomized trial published in Nature Medicine reports that an artificial intelligence-based clinician decision support system significantly improved diagnostic performance for suspected inherited retinal diseases, a group of conditions that are difficult to classify quickly and accurately in routine care.
The system, called Retina4IRD, was designed to help specialists predict likely genotype categories from retinal imaging. In the trial, specialists using the tool achieved substantially higher top-5 genetic prediction accuracy than specialists working without it. The result addresses a longstanding bottleneck in ophthalmology, where inherited retinal diseases often require intensive phenotyping, multidisciplinary interpretation and genetic testing before a firm diagnosis can be made.
Inherited retinal diseases, or IRDs, are not a single disorder but a collection of genetic conditions affecting the retina. Their clinical presentation can overlap, which means physicians often need to combine several kinds of evidence before narrowing down the likely cause. That process can be slow, resource-heavy and uneven across care settings, especially where deep subspecialty expertise is limited.
How Retina4IRD was built
According to the study, Retina4IRD predicts 17 genotype categories from retinal images. The research team developed the model using a Vision Transformer architecture pretrained with RETFound, then trained and validated it with multimodal imaging data drawn from genetically confirmed patients in China, South Korea and Poland.
The dataset included color fundus photographs and optical coherence tomography scans from 1,843 genetically confirmed patients, representing 3,376 eyes. That international training and validation set matters because IRD diagnosis can be affected by both disease diversity and differences in imaging practice across sites. External validation performance is therefore a key test of whether a system can travel beyond the institution where it was built.
In internal validation, the model reached a top-5 prediction accuracy of 0.904, with a 95% confidence interval of 0.896 to 0.912. In external validation, top-5 accuracy was 0.856, with a 95% confidence interval of 0.850 to 0.863. Those figures suggest the model retained strong performance even when evaluated outside its development environment, though with the expected drop from internal testing.
The study frames the tool not as a replacement for clinical judgment or genetic sequencing, but as a decision support system that can help clinicians prioritize the most likely genetic explanations earlier in the diagnostic pathway.
Randomized trial shows gains in specialist performance
The strongest evidence in the paper comes from a randomized controlled trial involving 300 participants with suspected IRD. Participants were assigned 1:1 either to a Retina4IRD-assisted specialist arm or to a specialist-only arm. Of those enrolled, 295 participants had available next-generation sequencing reports and were included in the final analysis.
The median age of those analyzed was 33 years, and 114 participants, or 38.6%, were female. The study’s primary endpoint was met. Specialists using the AI support tool achieved top-5 genetic accuracy of 88.5%, compared with 67.3% in the specialist-only group, a difference the paper reports as statistically significant with P less than 0.001.
The paper also states that secondary endpoints covering top-1 through top-4 accuracy all favored the Retina4IRD-assisted arm. Taken together, those results suggest the system did more than merely add extra possibilities to a long list. It improved ranking quality across the prediction stack, making the most plausible genetic categories more likely to appear near the top.
That distinction matters in practice. If a specialist can identify a smaller, higher-quality set of likely genotype categories from imaging, downstream testing and interpretation may become more focused. In settings where genetic testing capacity is constrained, a better front-end triage signal could help allocate specialist attention more efficiently.
What the findings could change
The paper’s significance lies in the type of evidence it provides. Many medical AI studies stop at retrospective validation, where algorithms are tested on historical data but not evaluated in live or quasi-live clinical workflows. Here, the researchers moved further by running a randomized comparison between AI-assisted and unassisted specialists.
That does not settle every deployment question. The supplied study text does not claim that Retina4IRD independently establishes a final diagnosis, nor does it suggest genetic testing can be skipped. Instead, the results point to a more practical near-term role for AI in rare-disease medicine: improving decision support in places where diagnosis currently depends on scarce expertise and time-intensive review.
The cross-border development dataset also hints at a broader ambition. IRDs are rare enough that assembling robust training data is difficult, and international collaboration is often necessary to capture clinically meaningful variation. If further studies confirm the results across wider populations and care systems, retina imaging models could become an increasingly important layer between image acquisition and molecular confirmation.
For now, the headline finding is straightforward. In a randomized, multicenter setting, an AI-based support system materially improved specialists’ ability to predict the genetic category behind suspected inherited retinal disease. In a field where diagnostic delay can shape treatment decisions, counseling and eligibility for future gene-targeted interventions, that is a consequential step.
This article is based on reporting by Nature Medicine. Read the original article.
Originally published on nature.com

