In the realm of knowledge-driven professions, a time may come when artificial intelligence surpasses human expertise. For the medical community, that moment seemed to arrive this past April, when a team of researchers from Harvard and Stanford shared findings from a study that compared the diagnostic capabilities of ChatGPT with those of hundreds of medical professionals. The results revealed that the AI chatbot outperformed its human counterparts in tackling complex medical cases, inciting mixed feelings within the medical community.
Lead author Adam Rodman expressed his reservations about the implications of such findings during a pre-publication press conference for the research featured in the journal Science. He emphasized that despite the thoroughness of the study, it did not definitively imply that ChatGPT or any AI system was ready for routine medical applications. Rodman’s caution mirrored sentiments shared by fellow experts, yet he acknowledged a widespread tendency for people to overlook such warnings. Alarmingly, AI is becoming integrated into the U.S. healthcare system, often sans adequate evidence or protective measures.
As I tuned into Rodman’s press conference, I received a notification from the medical institution where I practice as a pathologist. The message announced the availability of an “AI-powered clinical reasoning tool” for my use. This wasn’t the first such notification I had received; in fact, I've lost track of the numerous AI-based tools introduced in recent years, none of which have yet to receive FDA approval for medical use.
The rapid enthusiasm for AI in healthcare seems unprecedented. Traditionally, this field has been slow to embrace new technologies—despite continuing to rely on pagers and fax machines, which may perplex younger generations. This cautious approach is largely rooted in a commitment to safety, as even minor technological glitches could have severe consequences. Yet, there is an emerging trend encouraging healthcare professionals to experiment with the latest AI innovations, often guided by the vague caution that "AI can make mistakes."
These errors can yield serious consequences. Although Rodman's study highlights generative AI's potential to assist in diagnosing rare conditions and interpreting unusual symptoms, a separate randomized trial published in NEJM AI merely days before indicated that flawed outputs from AI models could mislead healthcare providers. Non-professionals may be equally misinformed; research from Oxford revealed that AI usage did not significantly enhance an individual's capacity to self-diagnose or assess others. A study from Mount Sinai suggested chatbots often fail to alert users about medical emergencies.
Misdiagnosis is only one of several issues. As AI technology seeps into healthcare, errors are emerging in unforeseen areas. After his press conference, Rodman recounted his surprise upon discovering that his hospital had begun using AI to draft messages to patients on his behalf, sometimes generating content he found to be “completely absurd.” A representative from Beth Israel Lahey Health noted that AI tool use is voluntary and includes appropriate training and support, with any generated content requiring physician approval.
One contributing factor to these problems is that health-related AI solutions can be implemented without oversight from the FDA. Tools identified as "clinical decision support" rather than medical devices typically circumvent the agency's scrutiny. For an AI application to qualify in this category, it generally must lean on pre-existing medical literature, refrain from analyzing medical images, elucidate its reasoning, and defer to physicians for diagnosis and treatment. Most of the generative AI tools available to doctors today fit these criteria.
Consumer health apps and devices may similarly evade FDA scrutiny, provided their purpose aligns with “promoting healthy lifestyles” rather than diagnosing specific conditions. Thus, Microsoft, OpenAI, Anthropic, and xAI caution users that their health-focused chatbots are not designed for medical care or diagnostic recommendations. However, the distinction can often be muddled. Elon Musk promotes his Grok chatbot for generating secondary medical opinions and interpreting X-ray and MRI images, while a promotional video for ChatGPT Health reassures users that their lab results fall within acceptable ranges and encourages ongoing cholesterol medication use.
Many of these applications also encourage users to link their medical records and health-monitoring devices. AI developers may collect this data without necessarily needing it for providing general health information. A new offering from the healthcare startup Hims & Hers, named Labs AI, assists users with interpretations from up to 130 biomarker tests, delivering thorough, personalized health assessments. As a physician who assesses lab results to offer tailored advice, I question how these apps differ from my practice.
Upon contacting the developers behind these tools, they insisted that they do not dispense actual medical advice. Microsoft’s AI health vice president clarified that the Copilot app offers “helpful information” for discussions with healthcare providers rather than providing definitive diagnoses. Similarly, Hims & Hers emphasized that Labs AI is designed to respect clinician roles in diagnosis and treatment recommendations. Anthropic and xAI did not respond to my attempts for commentary, while OpenAI declined to participate in this discussion.
Perhaps the line between doctor and algorithm is less distinct than it appears. One concept circulating in medical discourse proposes that AI tools should not merely be treated like standard medical devices. Given their ability to learn from data and customize their responses for individual patients, medical AIs might be more comparable to physicians than to tools like defibrillators—prompting suggestions that they should be evaluated similarly to doctors. Instead of requiring FDA approval for every capability, a chatbot might be assessed through a medical licensing examination and a supervised residency period.
For now, this concept remains largely theoretical. Haider Warraich, a cardiologist and program manager at the Advanced Research Projects Agency for Health, is spearheading efforts to gain traditional approvals for medical chatbots. His agency is financing the development of a heart-focused AI tool, which will undergo the full FDA evaluation process. Warraich hopes that this rigorous assessment will validate the chatbot's capacity to evaluate and treat patients independently. While Rodman supports this approach, he cautions that the process will take years, during which an influx of health AI tools could flood the market without thorough oversight.
The emergence of current AI health products is reminiscent of the rise of ride-sharing services like Uber and Lyft in the 2010s. The taxi industry, encumbered by stringent regulations, faced challenges from new entrants who bypassed or sidestepped these rules to quickly gather a loyal user base. Before long, regulatory bodies had to adapt their frameworks to accommodate the new reality. A similar scenario may unfold in healthcare. Will regulations designed to ensure safety and efficacy of medical products remain intact? Or will they be compromised or eliminated to accommodate the tools that have quickly gained popularity?
Answers will soon emerge. The healthcare landscape isn't poised to "wait for evidence to accumulate," as Rodman indicates. According to a 2026 survey by the American Medical Association, a staggering 80% of physicians are currently incorporating AI tools into their practice, and patient engagement is rapidly following suit. While the advantages of AI continue to be debated, their allure proves difficult to resist.



