Editorial illustration for Study: Medical AI's Helpfulness Depends on User's Own Expertise
Medical AI Helps Experts, Hurts Novices: MIT Study
Study: Medical AI's Helpfulness Depends on User's Own Expertise
Doctors and laypeople don't lean on artificial intelligence the same way, and a new study out of MIT suggests that difference could determine whether AI helps or hurts a diagnosis. Researchers tested non-experts and primary care providers on skin disease diagnosis, giving some of them AI assistance and, in certain cases, explanations of how the AI reached its conclusion. Those explanations came in different forms: heat maps flagging the image regions a model weighted most heavily, or plain-language reasoning generated by a large language model.
Both groups got more accurate with AI help. But the study found the reasons diverged sharply. Non-experts leaned on the AI's explanations regardless of whether the underlying prediction was correct, and rated vaguer, more generic explanations as more convincing.
Clinicians didn't fall for bad AI advice in the same way, and actually performed best when given a raw prediction with no explanation attached. The findings point to a design problem for anyone building diagnostic AI tools: an explanation feature that helps one user group can mislead another, depending on how much medical training that user already has.
A new study by researchers at MIT and elsewhere found that, while AI assistance generally improved the accuracy of non-experts and clinicians in diagnosing skin diseases, AI explainability methods had different impacts depending on the users’ knowledge level.
Why this matters
This study lands at an awkward moment for anyone building consumer-facing medical AI. The MIT team's finding that explainability features can actively mislead non-experts complicates the industry's default pitch: that showing users a model's reasoning automatically builds justified trust. It doesn't. Daneshjou's point about patients being "led astray" by erroneous explainable outputs should worry every team shipping symptom checkers or diagnostic assistants to the general public, not just hospitals with trained clinicians reading the same interface.
For developers and founders, the lesson is that explainability isn't a universal feature you bolt on and check off. It's a design choice that needs testing against the actual audience, whether that's a dermatologist or someone googling a mole at midnight. Researchers building these systems should be measuring outcomes by user expertise, not just aggregate accuracy gains, because a tool that helps specialists and hurts novices isn't net positive, it's a liability with good PR. Expect regulators and hospital systems to start asking for that breakdown soon.
Common Questions Answered
How does AI assistance impact diagnosis accuracy differently between non-experts and primary care providers?
According to the MIT study, AI assistance generally improved diagnostic accuracy for both non-experts and clinicians when diagnosing skin diseases, but the degree of improvement varied based on the user's existing knowledge level. The research found that users with different expertise levels responded differently to the same AI tools, suggesting that a one-size-fits-all approach to medical AI may not be effective across all user groups.
What did the MIT researchers discover about AI explainability methods and user expertise?
The MIT researchers found that AI explainability methods—such as heat maps highlighting image regions the model weighted most heavily or plain-language explanations—had different impacts depending on users' knowledge levels. Notably, explainability features could actively mislead non-experts rather than build justified trust, complicating the industry assumption that showing a model's reasoning automatically increases user confidence in the AI's conclusions.
Why does the MIT study's finding about explainability pose challenges for consumer-facing medical AI companies?
The study reveals that explainability features don't automatically build justified trust in medical AI systems, which contradicts the default pitch used by many companies developing consumer-facing diagnostic tools. The research suggests that non-experts can be 'led astray' by erroneous explainable outputs, meaning companies shipping symptom checkers and diagnostic assistants to the general public need to reconsider how they present AI reasoning to users without medical expertise.
What were the different forms of AI explanations tested in the skin disease diagnosis study?
The MIT researchers tested multiple types of AI explanations with study participants, including heat maps that flagged the image regions a model weighted most heavily and plain-language explanations of how the AI reached its diagnostic conclusion. These different explanation formats were provided to both non-experts and primary care providers to evaluate how various explainability approaches affected their diagnostic accuracy and trust in the AI system.
Further Reading
- People Overtrust AI-Generated Medical Advice despite Low Accuracy - MIT Media Lab / NEJM AI
- Reliability of LLMs as medical assistants for the general public: a randomized preregistered study - Nature Medicine
- Combining Human Expertise with Artificial Intelligence - MIT Economics
- Artificial intelligence in medicine: the influence of medical expertise and perceived causability on medical AI risk and benefit perception - Taylor & Francis
- Do as AI say: susceptibility in deployment of clinical decision-aids - Nature Digital Medicine