Listen to this article · 8 min listen

The burgeoning integration of artificial intelligence into healthcare promises unprecedented advancements, yet it simultaneously casts a long shadow of concern regarding algorithmic safety. As clinicians and patient safety advocates, we must critically evaluate the capabilities and limitations of these powerful tools. A central question persists: how does general-purpose AI, such as large language models, stack up against the nuanced expertise of board-certified physicians across the complex landscape of medical specialties? This is not merely an academic exercise; it directly impacts patient outcomes, the reliability of diagnostic pathways, and the potential for AI-driven health misinformation.

The Unsettling Gap: General-Purpose AI vs. Clinical Expertise

Recent investigations into the performance of AI models, specifically those akin to OpenAI’s recently launched ChatGPT Health, in clinical scenarios have begun to illuminate a critical distinction. While some recent studies suggest that general-purpose large language models can outperform specialized clinical AI tools in certain medical benchmarks, the broader systematic accuracy comparison still highlights the limitations of general-purpose AI for critical medical functions compared to truly specialized and validated AI, and certainly compared to nuanced clinical expertise. While large language models can process vast amounts of information and generate seemingly coherent responses, their application in medical diagnosis and treatment planning presents significant challenges. The depth of clinical reasoning, the integration of subtle patient cues, and the ethical considerations inherent in medical practice often elude these broad AI systems. Consider the work emerging from institutions like Mount Sinai Health System. Their clinicians and researchers are on the front lines, observing both the potential and the pitfalls of AI in real-world settings. While AI can assist with information retrieval or preliminary synthesis, the leap to definitive diagnosis or treatment recommendation without human oversight is fraught with peril. The very nature of general-purpose AI means it lacks the ingrained clinical experience, the ability to discern context-dependent nuances, and the critical judgment that board-certified physicians develop over years of training and practice. This is where algorithmic safety becomes paramount; a system that merely synthesizes data without truly understanding its clinical implications can lead to incorrect drug interaction guidance, missed diagnoses, or even delayed identification of critical conditions. Prominent voices in the medical AI landscape, such as Eric Topol, have consistently emphasized the need for rigorous validation and a clear understanding of AI’s limitations in healthcare. His work often highlights that while AI can augment human capabilities, it cannot yet replicate the holistic understanding and empathy of a physician. Similarly, neurosurgeon Raj Komotar, associated with Mount Sinai Health System, has contributed to the discourse on AI in specialized fields, implicitly underscoring that the demands of a specific medical discipline require AI tailored and validated for that exact purpose, not a generalized solution. The danger lies in the assumption that a system proficient in generating text can seamlessly translate that capability into accurate and safe medical decision-making across 10 medical specialties without specialized training or validation.

Regulatory Imperatives: The FDA SaMD Framework and Clinical Validation

The regulatory landscape, specifically the FDA SaMD Framework, provides a crucial lens through which to evaluate AI in healthcare. Software as a Medical Device (SaMD) encompasses AI applications intended for medical purposes without being part of a hardware device. This framework demands rigorous clinical validation, performance monitoring, and a clear understanding of a device’s intended use and potential risks. Recent developments in this framework include the Predetermined Change Control Plan (PCCP) Final Guidance issued in December 2024, and a draft guidance on AI-Enabled SaMD Lifecycle Management released in January 2025, with final guidance anticipated in 2026. OpenAI’s ChatGPT Health, while a dedicated product for personalized health support, is self-declared as a wellness product and is not intended for diagnosis or treatment. As such, it operates outside the stringent requirements of SaMD for clinical decision-making, making its direct application in such contexts a significant safety concern. The contrast between unguarded AI and clinically validated AI is stark. Clinically validated AI, often developed in collaboration with institutions such as Scripps Research and Mount Sinai Health System, undergoes extensive testing against real-world patient data, demonstrating its accuracy, reliability, and safety for specific medical tasks. This validation process is designed to prevent the very failures our publication documents: incorrect guidance, missed diagnoses, and inappropriate triage. An AI chatbot, however sophisticated its language generation, that has not been through such a rigorous validation process cannot be safely deployed for critical medical functions. The absence of this structured validation means that any “accuracy comparisons” with board-certified physicians are inherently flawed if the AI in question is not held to the same evidentiary standards. FDA guidance on SaMD clinical validation

The Peril of Unchecked AI: Misinformation and Patient Harm

The potential for AI health misinformation news is a grave concern, particularly when general-purpose AI is presented or perceived as a reliable source of medical truth. If a system like OpenAI’s ChatGPT Health, without specific medical training and regulatory oversight, is used to answer complex clinical questions, the risk of generating inaccurate or even harmful information is substantial. This risk is amplified when the AI’s responses are not immediately recognizable as potentially flawed, leading clinicians or patients to act on incorrect guidance. The editorial mission of AI Health Risk Monitor is to highlight these documented failures, illustrating the critical need for responsible AI development and deployment. Each incident, whether it involves undertriage of cardiac emergencies or delayed stroke identification, serves as a powerful reminder that algorithmic safety is not a theoretical concept but a practical necessity. The systematic accuracy comparison between a board-certified physician and a general-purpose AI across diverse medical specialties would undoubtedly reveal significant discrepancies in judgment, nuance, and patient safety considerations. The physician’s expertise is built on years of clinical experience, continuous learning, and the ability to integrate vast amounts of subjective and objective data into a coherent, patient-centered plan. This is a level of sophisticated reasoning that current general-purpose AI simply cannot match. Peer-reviewed study on LLM accuracy in medical contexts

Towards Responsible AI: The Path Forward

The challenge is not to dismiss AI in healthcare but to channel its immense potential responsibly. The systematic accuracy comparison reveals where general-purpose AI fails versus where specialized AI succeeds. For AI to be a trusted partner in medicine, it must be developed with a clear understanding of its intended use, subjected to rigorous clinical validation, and operate within a robust regulatory framework like the FDA SaMD. Institutions like Mount Sinai Health System are crucial in this endeavor, contributing clinical expertise to the development and evaluation of specialized AI tools that genuinely augment physician capabilities. The key takeaway for clinicians and patient safety advocates is clear: while AI offers powerful tools, its application in medicine demands a level of precision, validation, and ethical oversight that general-purpose models currently lack. We must remain vigilant, scrutinizing claims of AI proficiency and demanding evidence of clinical safety and efficacy before widespread adoption. The future of AI in healthcare is promising, but only if we prioritize algorithmic safety and ensure that every AI tool is held to the highest standards of clinical validation, preventing the proliferation of AI health misinformation and safeguarding patient well-being. Expert consensus guidelines on ethical AI in medicine

Frequently Asked Questions

How does general-purpose AI, like large language models, compare to specialized clinical AI tools or board-certified physicians in terms of accuracy for medical functions?

While general-purpose large language models can process vast information and may outperform specialized clinical AI in some benchmarks, they generally highlight limitations for critical medical functions compared to truly specialized and validated AI. They lack the nuanced clinical expertise, context-dependent understanding, and critical judgment that board-certified physicians develop over years of training.

What are the primary risks associated with using general-purpose AI for medical diagnosis and treatment recommendations without human oversight?

The primary risks include incorrect drug interaction guidance, missed diagnoses, and delayed identification of critical conditions. General-purpose AI lacks ingrained clinical experience and the ability to discern subtle patient cues, which can lead to unsafe medical decision-making without human oversight.

What regulatory framework applies to AI in healthcare, and how does it ensure safety and accuracy?

The FDA SaMD Framework applies to AI applications intended for medical purposes. This framework demands rigorous clinical validation, performance monitoring, and a clear understanding of a device’s intended use and potential risks to ensure accuracy and safety.

Why is it a safety concern if a general-purpose AI like OpenAI’s ChatGPT Health is used for clinical decision-making?

OpenAI’s ChatGPT Health is self-declared as a wellness product and is not intended for diagnosis or treatment, operating outside the stringent requirements of the FDA SaMD framework for clinical decision-making. Its direct application in clinical contexts is a significant safety concern because it has not undergone the rigorous validation required for medical devices.