A model designed for the pressures of the operating room
Intraoperative pathology sits at the center of precision surgery, giving surgeons the tissue-level information they need while a patient is still on the table. Yet the source of that information — the frozen section — has long been a bottleneck. The work is diagnostically complex, the samples are difficult to interpret under time pressure, and high-quality frozen-section datasets large enough to train modern computational tools have been scarce. Those constraints have limited how much clinical impact intraoperative pathology can realistically deliver.
Computational pathology has advanced considerably in recent years, but the field has faced a persistent credibility gap: a shortage of large-scale, prospective validation has kept most algorithms out of routine surgical workflows. A newly published study in Nature Medicine describes an effort to close that gap directly. Researchers introduced CRISP, a vision-based pathology foundation model built exclusively from frozen-section slides, with the explicit goal of providing Clinically-oriented Robust Intraoperative Support for Pathology.
The work appears as an open-access, accepted manuscript released early so that peer-reviewed findings can circulate faster. The version carries a permanent DOI and is citable, though the authors note it remains subject to further edits before being replaced by the final Version of Record.
Training on scale: 100,000 frozen sections from ten centers
The foundation model was developed using more than 100,000 frozen sections collected from ten medical centers. That dataset is notable not only for its size but for its composition. Frozen sections differ from the formalin-fixed, paraffin-embedded slides that dominate most computational pathology research — they are prepared rapidly, during surgery, and their visual characteristics reflect that urgency. By training exclusively on frozen-section material, the developers aimed to build a system matched to the actual conditions of intraoperative decision-making rather than adapting a model built for a different imaging context.
CRISP is described as clinically oriented, a framing that runs through both its training design and its evaluation strategy. Instead of optimizing purely for benchmark performance, the study treats real surgical utility as the measure of success.
Retrospective testing across nearly 100 diagnostic tasks
The evaluation phase spanned more than 15,000 intraoperative slides and close to 100 retrospective diagnostic tasks. Those tasks covered a deliberately broad range of clinical questions, including:
- Benign-versus-malignant discrimination
- Key intraoperative decision-making scenarios
- Pan-cancer detection
Across this battery of tasks, CRISP demonstrated robust generalization over six institutions, 14 tumor types and 24 anatomical sites. Importantly, that range included anatomical sites the model had not encountered during training as well as rare cancers — the kinds of edge cases where narrowly trained systems tend to fail. Generalizing to unseen sites and uncommon tumor types is a meaningful signal for a tool intended to support surgeons across varied specialties and hospital settings.
Prospective validation in more than 3,000 patients
The most consequential element of the study is its prospective component. CRISP was evaluated in a prospective cohort of over 3,000 patients, a setting that moves the assessment from curated retrospective slides into the messy conditions of real-world care. Under those conditions, the model sustained high diagnostic accuracy.
The headline clinical finding: CRISP directly informed surgical decisions in 92.6% of cases. That figure speaks to integration rather than mere prediction — the model's output was not simply compared against a reference label but was used in the moment when operative choices were being made.
Human-AI collaboration and workload relief
The study also examined what happens when pathologists and the model work together rather than in isolation. According to the reported results, human-AI collaboration produced several concrete benefits:
- Diagnostic workload fell by 35%
- 105 ancillary tests were avoided
- Detection of micrometastases reached 87.5% accuracy
Each of these outcomes addresses a different pressure point in surgical pathology. Reduced workload matters in departments where turnaround time is tight and staffing is limited. Avoided ancillary tests have implications for cost, resource use and the delay that additional testing can introduce during an operation. Improved micrometastasis detection targets one of the hardest problems in the field, where small deposits of tumor cells can be easy to miss and consequential to overlook.
Why frozen sections have resisted automation
Understanding the significance of CRISP requires appreciating why intraoperative pathology has lagged behind other areas of computational medicine. The diagnostic complexity is substantial: frozen-section morphology can be distorted by the freezing process, and pathologists must render judgments quickly with limited tissue. Meanwhile, the archival record of frozen sections at most institutions is fragmented compared with the vast libraries of permanent slides available for research.
That combination — hard images, fast decisions, thin data — is precisely the environment in which foundation models, which depend on large and representative training corpora, have struggled to establish themselves. Assembling more than 100,000 frozen sections across ten centers represents a coordinated, multi-institutional effort to overcome that structural obstacle.
What the results suggest for surgical practice
Taken together, the findings position CRISP as a clinically oriented approach to AI-driven intraoperative pathology, one with the potential to support surgical decision-making and to help translate computational methods into everyday practice. The study's design — multi-center training, broad retrospective testing including unseen sites and rare cancers, followed by prospective validation in thousands of patients — maps onto the evidentiary path that regulators and hospital systems typically expect before a tool is trusted inside an operating theater.
Questions inevitably remain. Sustained performance across a wider array of institutions, long-term monitoring for drift, and the practical logistics of deploying a foundation model in time-sensitive surgical settings are all areas that further study will need to address. The manuscript itself is an early, citable release rather than the final version of record, and the authors acknowledge that the content may be updated.
Still, the reported numbers give a concrete sense of what human-AI collaboration in the operating room might look like: a diagnostic partner that informed surgical decisions in the large majority of cases, absorbed roughly a third of the reporting burden, spared more than a hundred ancillary tests, and improved detection of micrometastases. For a specialty where minutes matter and accuracy is non-negotiable, that is a meaningful step toward making computational pathology a routine part of the surgical workflow rather than a promising research prospect.
This article is based on reporting by Nature Medicine. Read the original article.
Originally published on nature.com






