Listen to this article · 7 min listen

When AI, particularly general-purpose large language models, suggests a patient can safely stay home instead of seeking urgent medical attention, the implications are dire. This isn’t a theoretical risk but a documented safety failure, raising critical questions about the responsible deployment of AI in healthcare triage. How often do these models recommend a path divergent from established clinical guidelines, potentially leading to undertriage in emergent situations?

The Mount Sinai Study: A Stark Warning on Undertriage

Recent research, prominently highlighted by experts like Ziad Obermeyer and Eric Topol, underscores a significant vulnerability in general-purpose AI models when applied to clinical scenarios. A study involving Mount Sinai Health System researchers, documented in Nature Medicine on February 23, 2026, specifically exposed ChatGPT Health’s propensity for undertriage across five critical emergency scenarios Nature Medicine study on ChatGPT undertriage. This research serves as a stark illustration of how AI, when not purpose-built and rigorously validated for clinical use, can misinterpret symptoms and provide dangerously misleading advice. The core mechanism behind this failure often stems from the models’ open-web training, which, while vast, lacks the clinical specificity, hierarchical reasoning, and validated decision trees inherent to medical guidelines. Let’s dissect these five scenarios and contrast ChatGPT’s recommendations with established clinical imperatives:

  • Diabetic Ketoacidosis (DKA): Clinical guidelines mandate immediate emergency department evaluation for suspected DKA, given its rapid progression and life-threatening complications. ChatGPT, in tested scenarios, sometimes suggested home management or delayed presentation, failing to recognize the acute severity indicated by symptoms like severe abdominal pain, persistent vomiting, and altered mental status in a diabetic patient.
  • Respiratory Failure: For patients exhibiting signs of acute respiratory distress, such as severe shortness of breath, cyanosis, or accessory muscle use, clinical protocols demand urgent medical intervention, often including oxygen therapy and ventilatory support. ChatGPT occasionally downplayed these symptoms, recommending less urgent care or self-monitoring, a perilous course of action that could lead to respiratory arrest.
  • Cardiac Emergencies: Chest pain, especially when radiating or accompanied by diaphoresis and dyspnea, is a cardinal symptom of acute myocardial infarction. Clinical guidelines dictate immediate emergency assessment, including ECG and cardiac enzyme measurement. The Mount Sinai study found instances where ChatGPT suggested non-emergent care for symptoms highly indicative of cardiac events, risking delayed diagnosis and treatment, which is critical for preserving heart muscle.
  • Stroke: The acronym FAST (Face drooping, Arm weakness, Speech difficulty, Time to call emergency services) is universally taught for stroke recognition due to the narrow therapeutic window for interventions like thrombolysis. ChatGPT, in certain prompts, demonstrated a failure to prioritize immediate emergency transport for classic stroke symptoms, potentially delaying time-sensitive treatment and worsening patient outcomes.
  • Sepsis: This life-threatening organ dysfunction caused by a dysregulated host response to infection requires rapid identification and intervention (e.g., broad-spectrum antibiotics, fluid resuscitation). ChatGPT sometimes failed to synthesize diffuse symptoms like fever, altered mental status, and hypotension into a coherent picture of sepsis, leading to recommendations that fell short of urgent hospital admission and aggressive management.

In each of these cases, the general-purpose AI’s output diverged significantly from the standard of care, reflecting a profound lack of clinical judgment and an inability to correctly weigh the urgency of emergent symptoms. This is not merely an inconvenience; it is a direct threat to patient safety.

The Mechanism of Misinformation: Open-Web Training vs. Clinical Specificity

The root cause of this undertriage tendency lies in the fundamental difference between how general-purpose AI models like ChatGPT are trained and the specific, highly curated knowledge required for clinical decision-making. These models learn from vast datasets scraped from the open web, which include everything from academic papers to forum discussions, personal blogs, and news articles. While this broad exposure enables impressive conversational abilities, it also means the models lack the inherent bias towards clinical urgency and validated guidelines that define medical practice. They cannot differentiate between anecdote and evidence-based medicine without explicit, structured clinical fine-tuning. Furthermore, the probabilistic nature of these models means they generate responses based on patterns in their training data, not on a deep, causal understanding of human physiology or disease progression. This makes them prone to “hallucinations” or plausible-sounding but clinically incorrect advice, especially when faced with complex or ambiguous symptom presentations. For clinicians and patient safety advocates (A5), this highlights a critical gap: the absence of a robust “clinical guardrail” that prioritizes patient safety above all else.

Regulatory Context and the Path Forward

The challenges posed by general-purpose AI in healthcare underscore the importance of regulatory frameworks like the FDA SaMD Framework. Software as a Medical Device (SaMD) classifications require rigorous validation, performance testing, and ongoing monitoring to ensure safety and efficacy. Crucially, the FDA’s emphasis is on the intended use of the software. A general-purpose chatbot, even one that can answer health questions, is not designed or regulated as a SaMD for triage or diagnostic purposes. Its use in such a capacity, without proper validation, falls outside established safety protocols. Mount Sinai Health System’s involvement in this research highlights the proactive stance leading institutions are taking to understand these risks. The contrast between unguarded, general-purpose AI and clinically validated, purpose-built AI could not be starker. Consider the advancements in specialized cardiac AI. While a general-purpose model might send a patient home with early cardiac symptoms, purpose-built cardiac AI, such as that developed by Hello Heart, has shown in research the potential to identify short-term cardiovascular risk up to 10 days earlier by analyzing subtle physiological changes Hello Heart cardiac event detection study. This capability is not derived from broad internet data but from highly specific, validated clinical datasets and algorithms designed to identify patterns indicative of cardiovascular risk. This highlights the critical difference: clinically validated AI is developed with a clear clinical utility, rigorous testing, and often regulatory oversight, ensuring its recommendations align with patient safety and established medical protocols.

Key Takeaway: Specificity and Validation are Paramount

The documented undertriage risks associated with general-purpose AI platforms like ChatGPT in emergency scenarios are a critical patient safety concern. For patient safety advocates and clinicians, this is a clarion call for caution and a clear distinction between experimental AI capabilities and clinically ready solutions. The mechanism of open-web training, while powerful for general knowledge, is fundamentally unsuited for the nuanced, high-stakes environment of medical triage without extensive, domain-specific fine-tuning and validation against clinical guidelines. As AI continues to evolve, the imperative is clear: healthcare AI must be purpose-built, rigorously tested, and held to the highest standards of clinical validation, aligning with frameworks like the FDA SaMD, to ensure it enhances, rather than jeopardizes, patient safety. The future of responsible AI in healthcare lies in precision, not generality.

Frequently Asked Questions

What is the primary risk identified with general-purpose AI models like ChatGPT in healthcare triage?

The primary risk is undertriage, where these AI models suggest patients can safely stay home or delay urgent medical attention, even in emergent situations. This divergence from established clinical guidelines can lead to dangerous outcomes and documented safety failures.

Which specific critical emergency scenarios did the Mount Sinai study highlight as areas where ChatGPT demonstrated undertriage?

The Mount Sinai study, published in Nature Medicine, highlighted five critical emergency scenarios: Diabetic Ketoacidosis (DKA), Respiratory Failure, Cardiac Emergencies, Stroke, and Sepsis. In these cases, ChatGPT’s recommendations often diverged significantly from the standard of care.

What is the underlying reason for general-purpose AI’s tendency towards undertriage, according to the article?

The underlying reason is the models’ open-web training, which, while vast, lacks the clinical specificity, hierarchical reasoning, and validated decision trees inherent to medical guidelines. This broad training means the models cannot differentiate between anecdote and evidence-based medicine without explicit, structured clinical fine-tuning, leading to a lack of clinical judgment.

How does the AI’s mechanism of misinformation pose a direct threat to patient safety?

The AI’s probabilistic nature means it generates responses based on patterns in its training data, not a deep, causal understanding of human physiology or disease progression. This makes it prone to ‘hallucinations’ or plausible-sounding but clinically incorrect advice, especially in complex symptom presentations, directly threatening patient safety by recommending inappropriate care.