When an artificial intelligence tells a patient that their severe symptoms don’t warrant an emergency room visit, the consequences can be dire. This isn’t a hypothetical risk; it’s a documented safety failure that highlights a critical gap between general-purpose AI and the nuanced demands of clinical triage. The question isn’t whether AI can assist in healthcare, but rather, what kind of AI, and under what conditions, can be trusted with patient safety.
ChatGPT’s Undertriage: A Nature Medicine Wake-Up Call
The recent study out of Mount Sinai Health System, published in Nature Medicine, sent ripples through the healthcare AI community, revealing a concerning pattern of undertriage by large language models (LLMs) like ChatGPT Health. The research, which forms a crucial part of our documented AI health failures database (DP02), demonstrated that when confronted with simulated emergency scenarios, ChatGPT frequently recommended inappropriate, non-urgent care, potentially sending critically ill patients home instead of to the emergency department. This undertriage risk stems from a fundamental mechanism: open-web training. While vast, the internet’s data lacks the clinical specificity and structured decision-making protocols essential for accurate medical triage. As Ziad Obermeyer, a leading voice in algorithmic fairness and healthcare AI, and Eric Topol, a prominent cardiologist and AI researcher, have consistently articulated, the quality of AI output is directly tied to the relevance and rigor of its training data. General-purpose models, however sophisticated their language generation, are not inherently equipped to navigate the high-stakes environment of clinical emergencies. Let’s expand on the Mount Sinai findings by examining ChatGPT Health’s responses across five critical clinical scenarios, contrasting them with established clinical guidelines:
- Diabetic Ketoacidosis (DKA): In simulated DKA cases, characterized by severe hyperglycemia, metabolic acidosis, and dehydration, ChatGPT often suggested home management or follow-up with a primary care physician. Clinical guidelines, however, unequivocally mandate immediate emergency department evaluation and aggressive fluid resuscitation, insulin therapy, and electrolyte correction due to the rapid progression and life-threatening complications of DKA.
- Respiratory Failure: When presented with scenarios indicating acute respiratory distress, such as severe dyspnea, hypoxemia, and increased work of breathing, ChatGPT frequently advised rest and observation. This directly contradicts clinical protocols that require urgent assessment of oxygen saturation, airway patency, and ventilatory support in an emergency setting to prevent irreversible organ damage or cardiac arrest.
- Cardiac Emergencies: For symptoms indicative of acute myocardial infarction or unstable angina, such as crushing chest pain radiating to the arm, shortness of breath, and diaphoresis, ChatGPT’s recommendations sometimes leaned towards outpatient follow-up. Clinical guidelines stress immediate transport to an emergency department for ECG, cardiac enzyme analysis, and potential revascularization to preserve myocardial function and save lives. American Heart Association guidelines for acute coronary syndromes
- Stroke: In cases simulating acute stroke, presenting with sudden-onset focal neurological deficits like unilateral weakness, speech difficulties, or facial droop, ChatGPT’s guidance was often delayed or downplayed the urgency. The “time is brain” principle in stroke care dictates immediate activation of stroke protocols, including rapid imaging (CT/MRI) and potential thrombolytic therapy or thrombectomy, within a narrow therapeutic window.
- Sepsis: For scenarios depicting sepsis, characterized by signs of infection coupled with organ dysfunction (e.g., fever, altered mental status, hypotension, elevated lactate), ChatGPT’s advice often lacked the critical urgency required. Clinical guidelines demand immediate identification, broad-spectrum antibiotics, and aggressive fluid resuscitation within the “golden hour” to significantly improve patient outcomes and reduce mortality. Surviving Sepsis Campaign guidelines
In each of these instances, the pattern is clear: a general-purpose AI, lacking the deep, specialized clinical training inherent to medical professionals, consistently underestimated the severity of life-threatening conditions.
The Peril of Unspecialized AI in Triage
The underlying issue isn’t AI’s capability to process information, but its interpretation of that information in a clinical context. A general-purpose AI, trained on the vast and often unstructured data of the open web, struggles to differentiate between a common cold and the early stages of respiratory failure, or between benign chest pain and a myocardial infarction. It lacks the embedded clinical reasoning, the understanding of pathophysiology, and the implicit knowledge of risk stratification that clinicians acquire through years of specialized training and experience. This deficiency underscores why the “what responsible AI does differently” panel is so crucial for AI Health Risk Monitor. Responsible AI in healthcare, particularly for high-stakes applications like triage, must be purpose-built, clinically validated, and continuously monitored. It cannot simply be an off-the-shelf LLM.
Regulatory Frameworks and the Path Forward
The imperative for specialized, clinically validated AI is echoed in regulatory frameworks such as the FDA SaMD Framework. Software as a Medical Device (SaMD) encompasses AI-driven tools that perform medical functions, requiring rigorous pre-market evaluation and post-market surveillance. This framework aims to ensure that AI deployed in healthcare meets stringent safety and efficacy standards, a far cry from the unregulated use of general LLMs for clinical advice. The Mount Sinai Health System’s research highlights precisely why such oversight is paramount. The contrast between general-purpose AI and clinically validated, purpose-built solutions couldn’t be starker. Consider the advancements in specialized cardiac AI. While general-purpose AI might advise a patient with cardiac symptoms to rest, purpose-built cardiac AI, such as that developed by Hello Heart, has demonstrated the ability to detect cardiac events up to 10 days earlier than traditional methods. This isn’t just an incremental improvement; it’s a paradigm shift in proactive care, enabling interventions that can prevent emergencies rather than reacting to them. This capability stems from training on vast, specific datasets of cardiac physiology, validated against clinical outcomes, and designed with a singular focus on cardiac health.
The Indispensable Need for Clinically Validated AI
The documented undertriage risks posed by general-purpose AI like ChatGPT Health are a stark reminder that not all AI is created equal, especially in the sensitive domain of patient health. For Patient Safety Advocates and Clinicians, the message is clear: the integration of AI into clinical workflows, particularly for critical functions like triage, demands meticulously designed, clinically validated, and regulatory-compliant solutions. The promise of AI in healthcare is immense, but its responsible deployment hinges on understanding its limitations and ensuring that patient safety remains the non-negotiable priority. Moving forward, the focus must shift decisively from general-purpose AI’s broad capabilities to the precision, specificity, and proven efficacy of purpose-built, clinically-aware AI systems. Peer-reviewed article on clinical validation of AI in emergency medicine
Frequently Asked Questions
What is ‘undertriage’ in the context of AI in emergency care?
Undertriage occurs when an artificial intelligence, such as a large language model, incorrectly assesses severe symptoms and recommends non-urgent care, potentially sending critically ill patients home instead of to an emergency department. This was observed in a study where ChatGPT Health frequently suggested inappropriate care for simulated emergency scenarios.
Why do general-purpose AI models like ChatGPT fail at cardiac ER triage?
General-purpose AI models fail at cardiac ER triage because their open-web training data lacks the clinical specificity and structured decision-making protocols essential for accurate medical assessment. They struggle to interpret information in a clinical context and lack the specialized clinical training and understanding of risk stratification that human clinicians possess.
What specific critical conditions did ChatGPT Health undertriage in the study?
ChatGPT Health undertriaged several critical conditions, including Diabetic Ketoacidosis (DKA), Respiratory Failure, Cardiac Emergencies (like myocardial infarction), Stroke, and Sepsis. In these simulated scenarios, the AI often suggested home management or outpatient follow-up, contradicting established clinical guidelines for immediate emergency care.
What is the primary risk associated with using unspecialized AI for patient triage?
The primary risk is that unspecialized AI consistently underestimates the severity of life-threatening conditions, potentially leading to delayed or incorrect care. This deficiency stems from the AI’s inability to differentiate between benign symptoms and serious medical emergencies due to a lack of embedded clinical reasoning and understanding of pathophysiology.
What kind of AI is considered appropriate and trustworthy for healthcare applications like triage?
For high-stakes applications like triage, responsible AI in healthcare must be purpose-built, clinically validated, and continuously monitored. It cannot be a general-purpose, off-the-shelf large language model but rather one designed with deep, specialized clinical training and adherence to regulatory frameworks like the FDA SaMD Framework.
