Clinical prediction models need more than a published result
Clinical prediction models are increasingly used to support diagnostic and prognostic decisions. Their reported accuracy can influence how clinicians estimate risk, prioritize tests or discuss likely outcomes with patients. But a new analysis in Nature Medicine argues that one basic condition for evaluating such models remains uncommon: access to the analytical code behind them.
The study reviewed open-access research articles that cited the TRIPOD or TRIPOD+AI reporting frameworks, which concern transparent reporting of prediction-model studies. Researchers used a large-language-model-assisted pipeline to screen the literature, identify repository links and evaluate the repositories they could retrieve using 14 predefined reproducibility-related features.
Of 3,967 articles examined, 482 included a statement saying that code was shared. That is 12.2% of the sample. The finding does not establish that every remaining study lacked accessible code, but it does show that explicit code-sharing statements were still the exception in the reviewed literature.
Availability is not the same as reproducibility
The researchers’ central point is that a repository link alone does not make an analysis easy to check or reuse. They found substantial variation in the reproducibility-related features of the repositories they assessed. In practice, another researcher may need more than source files to understand whether a model can be recreated: documentation, information about software dependencies and an executable project structure can all matter.
That distinction is particularly important for clinical prediction work. A model may be developed from a specific cohort, preprocessing workflow and set of outcome definitions. If the analytical steps cannot be independently inspected, it is harder for outside researchers to assess how choices in the workflow affected reported performance, or whether the work can be adapted and tested in a different setting.
The review also found that code-sharing prevalence differed widely by journal and country. The supplied study text does not identify which journals or countries led or lagged, but the variation suggests that norms and editorial expectations are influencing whether researchers make their work available.
An AI-assisted review of research practice
The authors used a large-language-model-assisted process in a methodological role: screening the relevant articles, extracting repository links and helping assess repositories against a predefined feature set. The approach allowed the team to examine thousands of papers, while the substantive measure remained focused on concrete reporting and repository characteristics.
Its use also reflects a wider challenge in research governance. As publication volumes grow, manual audits of transparency practices become difficult to scale. Structured, carefully designed automated pipelines can help map those practices, provided their outputs are evaluated against clearly defined criteria. In this case, the criteria were tied to code availability and features relevant to reproducibility rather than to an automated judgment about the scientific quality of an individual model.
Why the gap matters for deployment
Clinical models are intended to inform real-world decisions, so their reliability and generalizability need close scrutiny. A model that appears effective in a published paper may perform differently when implemented elsewhere, where patient populations, data collection and clinical workflows differ. Independent assessment of the analytical process is one way to examine the foundations of those claims.
Open code cannot solve every reproducibility problem. Patient data can be restricted for privacy and legal reasons, and external users may not be able to recreate an analysis exactly without access to the original data. Even so, clear analytical code, dependency details and usable documentation can make the work more inspectable and can help researchers understand what would be required to test it responsibly with another dataset.
The authors frame the results as evidence for clearer expectations beyond a simple declaration that code exists. They say that documentation, dependency specification and executable structure should be part of stronger practice. Those elements can reduce ambiguity for readers and for teams attempting to reproduce or extend a study.
TRIPOD-Code is the next policy target
The review was conducted to inform TRIPOD-Code, a reporting guideline focused on code availability and reproducibility. Reporting guidelines do not themselves guarantee that researchers will release complete, usable projects. They can, however, make expectations visible to authors, reviewers and journals, and create a shared vocabulary for identifying what has been provided.
The result is a concrete benchmark: in this literature sample, only about one in eight papers included a code-sharing statement. For a field in which models are moving toward clinical use, the authors argue that improving the quality and clarity of shared analytical work could make research more usable and help support safer translation into practice.
The study therefore shifts attention from whether a paper merely mentions code to whether an independent reader could understand and run the analytical workflow. That is a higher bar, but one closely aligned with the evidence needed to judge clinical prediction models before they are relied on in care settings.
This article is based on reporting by Nature Medicine. Read the original article.
Originally published on nature.com







