Explainable AI did not affect everyone the same way
A new study in Nature Medicine reports that explainable AI can improve dermatological diagnosis support, but its benefits depend heavily on who is using it and when the AI’s answer is shown. In two large experiments involving 623 lay people and 153 primary care physicians, researchers found that assistance from a fairness-constrained AI model improved overall diagnostic accuracy and reduced performance disparities across skin tones. But the explanatory layer built on top of that system produced sharply different effects for non-experts and clinicians.
The central result is less a simple endorsement of explainable AI than a warning about its uneven influence. Lay users became more vulnerable to automation bias: when the AI diagnosis was correct, their performance improved, but when the system was wrong, their performance fell. Primary care physicians, by contrast, proved more resilient. The study says they benefited regardless of whether the AI’s diagnosis was right or wrong, suggesting that experience changed how explanations were interpreted and weighed.
That finding matters because medical AI is increasingly being aimed at two very different audiences at once. One is clinicians using software in professional settings. The other is the public, which is now encountering AI-powered symptom checkers, image interpreters, and large language model interfaces directly. This study suggests those two contexts should not be treated as interchangeable.
A fairness-constrained model improved baseline performance
The researchers built the experiments around a dermatology AI system trained with fairness constraints intended to balance performance across skin tones. That design choice addressed one of the most persistent concerns in medical imaging AI: systems can perform unevenly across demographic groups if their training data or optimization process leaves some populations underrepresented.
According to the study abstract, the fairness-constrained model improved final diagnostic accuracy for both lay users and primary care physicians. It also reduced skin-tone-related performance disparities. That is an important baseline result because it shows that model design choices can affect not only average performance but also equity in how assistance is distributed across users and cases.
In practical terms, the model was not simply providing raw predictions. It was part of a human-AI decision workflow in which users saw AI outputs and, in some conditions, multimodal large language model explanations. The researchers were therefore able to study not only whether the AI was accurate, but also how its presentation changed human judgment.
Explanations increased risk for lay users
Explainable AI is often presented as a remedy for black-box decision-making. If users can see why a system produced an answer, the thinking goes, they will trust it more appropriately and correct it when it fails. This study adds to a growing body of evidence that the relationship is more complicated.
For lay people, the explanations appear to have amplified the tendency to follow the machine. When the system’s diagnosis was correct, that increased reliance helped. When the system made an error, it hurt. The study describes this pattern as stronger automation bias, meaning the explanatory interface may have made the model’s conclusions feel more authoritative even when they should have been questioned.
That distinction is crucial for consumer-facing health tools. A polished explanation can create an impression of transparency without guaranteeing that the underlying recommendation is reliable in the specific case a user is looking at. For a non-expert, the extra detail may not create true understanding. It may instead create confidence.
The researchers characterize multimodal large language model explanations as a potential double-edged sword. In this context, that phrase is not rhetorical. It captures the core design problem: the same interface feature that helps one user make sense of a recommendation may push another user toward over-trust.
Clinicians responded differently
The primary care physicians in the study did not show the same vulnerability. The abstract says they remained resilient and benefited irrespective of whether the AI model’s diagnosis was accurate. That does not mean clinicians were immune to influence, but it does suggest they interacted with the system from a different cognitive starting point.
Professional training likely gave physicians a stronger internal frame for evaluating the model’s output against their own reasoning. Instead of accepting the explanation as a substitute for judgment, they may have used it as one more input in a broader diagnostic process. The study abstract does not spell out every mechanism behind this difference, but the contrast is one of its most policy-relevant conclusions.
Healthcare systems often discuss AI deployment as though one interface can serve many roles. This research points in the opposite direction. Tools intended for clinicians may need different explanation styles, confidence cues, and sequencing rules than tools intended for the public. An explanation that is useful in one setting may be actively misleading in another.
Timing mattered too
The study also found that presenting the AI diagnosis before human decision-making may create stronger anchoring bias. Anchoring happens when an early piece of information disproportionately shapes later judgment. In a diagnostic workflow, that could mean users become locked onto the AI’s first suggestion and insufficiently consider alternatives.
This result is especially relevant for software design. It means the safety question is not only whether to include explanations, but also when to reveal predictions and how to structure the sequence of interaction. A system that asks a clinician or patient for an initial assessment before showing the AI’s answer may produce different outcomes from one that leads with the machine recommendation.
That kind of interface choice can look minor from a product perspective, but it may have outsized effects on decision quality. The paper therefore contributes to a broader shift in medical AI evaluation: performance metrics alone are not enough. Human factors, workflow order, and explanation framing can materially change the real-world result.
Implications for medical AI deployment
The study does not argue that explainable AI should be abandoned. Its evidence instead supports a more precise claim: explanation features must be tested in the context of actual users, actual tasks, and actual decision sequences. In dermatology support, fairness-aware training improved outcomes, but explainability introduced new behavioral risks, particularly for non-experts.
That matters for regulators, hospitals, and consumer health developers alike. A tool that appears safer because it explains itself may still create avoidable harm if those explanations increase automation bias or anchoring. Conversely, carefully designed AI assistance can improve accuracy and reduce disparities if the model and interface are aligned with the user’s expertise.
As large language models move deeper into medical workflows, this paper offers a timely reminder that transparency is not automatically protective. In some cases it helps; in others it can distort judgment. The next phase of medical AI design will likely depend less on adding explanations everywhere than on identifying which explanations, for which users, and at what moment in the decision process, actually improve care.
This article is based on reporting by Nature Medicine. Read the original article.
Originally published on nature.com

