Artists learn a trick they use when planning a composition: if you squint, you eliminate the details and shades of gray; the world becomes black and white, light and shadow.
But life isn’t black and white; life is shades of gray.
Two recent articles offer very different views of AI in clinical care. They both acknowledge that AI outperforms physicians on standardized knowledge tests and in solving some difficult diagnostic challenges, however from my perspective each article portrays the role of AI in clinical medicine too much as either bad or good.
Rosenthal and Verghese (The iPatient Meets the iDoctor) see AI as, “creating an ‘iDoctor,’ an uber-representation of ourselves as physicians,” and view AI – even with its benefits – as “further hollowing out of longitudinal patient-physician relationships”. The authors highlight the sanctity of the ritual of a careful physical examination performed by a “living, breathing physician” and worry that “a further decline in or even total absence of physical examinations seems likely with increasing AI integration into practice…” And I suspect that would be true… in a black-and-white world.
By contrast, Emanuel, et al. (Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care?) disagree with the “prevailing view” that AI should simply support clinicians since in many aspects of clinical care, AI is better than humans and, “when AI-alone performance is consistently superior to human-alone performance, AI alone surpasses human-AI hybrids”. And I suspect that would be true… in a black-and-white world.
But life isn’t black and white; life is shades of gray.
A recent article in Nature Medicine (General-Purpose Large Language Models Outperform Specialized Clinical AI Tools on Medical Benchmarks) caused quite a stir when it reported that frontier large language models (LLMs) outperform clinical AI tools in several tests of medical knowledge. However, clinical care is much more than retrieving randomized trials and guidelines because medicine is about as gray as it gets. Many of the conversations I have with patients are about subtleties and the potential benefits and harms of diagnostic tests and treatments for them as individuals. Few decisions are black and white. For example, a typical question I’m asked as a cardiologist is, “How often should a person get a coronary calcium scan?” While a question like this might be on a multiple-choice test, it shouldn’t be. There’s not a single right answer. It’s gray.
So are frontier LLMs able to see shades of gray? One frontier AI model responds that coronary calcium scans should be obtained, “every 3 to 5 years” and another offers “typical intervals” for scanning. By contrast, Dyna AI provides a far more nuanced response. It references recommendations from specialty societies and notes potential harms of coronary calcium scanning, something that is not even mentioned by the frontier LLMs I queried.
Screenshot: Dyna AI response to query
I can understand why clinicians who are uncomfortable with the grays would find the black-and-white answers from frontier models of AI comforting. Squinting is something some clinicians have learned to do, not as an artist’s trick, but as a way to deal with the grays of day-to-day clinical medicine.
But in a world of grays, I say “Open your eyes!” Clinicians supported by an AI-enabled clinical decision support tool that embraces the grays like Dyna AI make for human-AI hybrids I’ll bet on any day.