Clinicians did best with no explanation at all

Research

Clinicians did best with no explanation at all

Explainable AI is supposed to help people judge a model's output. In skin-disease diagnosis it did the opposite for the people best qualified to judge.

A dim hospital corridor at night with a single lit doorway at the far end

Published

July 18, 2026

Reading time

2 minutes

Perspective

Research

Topics

healthcare · evaluation · mit

A study published in Nature Medicine in August 2026 tested non-experts and primary care physicians diagnosing skin diseases, with and without several kinds of explainable-AI support — heat maps, LLM explanations, similar-image comparison, and predictions alone.

The result splits cleanly by expertise, and not in the direction the field usually assumes.

The finding

Non-experts improved — but through deference. They trusted LLM explanations largely regardless of whether those explanations were accurate.

Clinicians performed best receiving only predictions, with no explanation at all. They were, in the study's terms, "resilient to incorrect AI explanations."

So explanations helped the group that could not evaluate them, and were unnecessary — at best — for the group that could.

The detail that should worry people

Two observations from the study:

Non-experts "found explanations more convincing when they were vague or generic."

And they were "more confident about their wrong answers when aided by an LLM."

Read together, that is a description of a persuasion mechanism rather than a comprehension aid. Vagueness reading as credibility is the opposite of what an explanation is for, and confidence rising on wrong answers is the specific failure that makes a decision-support tool dangerous rather than merely unhelpful.

What this does to the case for explainable AI

The standard argument is that explanations let a human audit the model, catching errors a bare prediction would smuggle through.

This study suggests auditing requires expertise the explanation cannot supply. If you already know enough to evaluate the reasoning, you did not need it. If you do not, the explanation persuades you instead of informing you.

That is not an argument against explainability. It is an argument that "add explanations" is not automatically a safety improvement, and can invert into a harm.

Where the design conclusion lands

For expert users, predictions alone performed best — which implies simpler interfaces, not richer ones.

For non-experts, the problem is calibration rather than explanation. A system that conveys uncertainty honestly gives a novice something actionable. One that produces fluent justification for a wrong answer does not, and this study measured exactly how that goes.

Why it generalises beyond dermatology

The structure is common: a tool proposes, a human disposes, and disposing well requires the judgement the tool was meant to substitute for. Code review, grading, literature triage all have this shape.

Anyone reporting an average accuracy gain from AI assistance without stratifying by user expertise is reporting a number that could be concealing this exact reversal.

Study led by Orson Xu (Columbia) with Marzyeh Ghassemi (MIT) and Roxana Daneshjou (Stanford). Source: MIT News

Continue reading

More from COREXA