Generative AI in healthcare is getting a lot of attention for its potential to fix administrative headaches, and for good reason. But the jump from optimizing back-office paperwork to running high-stakes clinical functions like emergency department triage is huge, and the challenges are being seriously underestimated. For any venture capital firm looking at this space, you have to understand why general-purpose large language models (LLMs) are falling short in these life-or-death scenarios if you want to de-risk your investments and back responsible tech.
Administrative Use vs. Clinical Reality
You can see why LLMs are so appealing in healthcare. They’re great at chewing through mountains of unstructured data, which makes them a good fit for summarizing medical records, handling prior authorizations, or even drafting first-pass clinical notes. We’re already seeing companies like Epic Systems, working with partners like Microsoft Azure AI, build these features to lighten the load on clinicians. But we can’t let the success of LLMs in these admin-support roles fool us into thinking they’re ready for direct clinical decision-making. Especially not in a place like the ER. Emergency department triage isn’t just about processing information. It requires a kind of nuanced clinical judgment and pattern recognition that comes from years of experience, plus a feel for contextual patient factors that are often buried or missing in the initial intake data. At its heart, the problem is that LLMs are built for pattern matching on text, while expert human practice is about inferential reasoning, causality, and managing diagnostic uncertainty. The American Medical Association (AMA) keeps flagging these clinical safety and efficacy issues, rightly insisting on serious validation before these models get anywhere near a patient AMA position on AI clinical safety.
Where Triage Accuracy Breaks Down: The Data on LLM Failures
The peer-reviewed research is pretty clear: LLMs are struggling with clinical triage. Study after study in journals like JAMA Network Open has put these models up against human nurses in simulated triage scenarios, and the results show a real drop in accuracy. They’re particularly bad at spotting high-acuity cases or telling the difference between conditions that look similar at first but have wildly different outcomes. For example, one recent analysis of LLM performance on ER chief complaints found that while the models did fine with common, low-stakes problems, their sensitivity and specificity tanked for less common but critical conditions, which could easily lead to undertriage JAMA Network Open LLM triage study. This isn’t just an academic problem. It’s a real safety risk. An LLM-powered system that labels a potential heart attack as a non-urgent stomach ache is going to get someone killed. The fundamental challenge is that general-purpose LLMs work by learning statistical relationships from huge, often messy datasets, which gives them impressive conversational ability but no actual clinical knowledge or understanding of physiology. They don’t have the medical “common sense” that lets a nurse or doctor spot subtle red flags and pull together disparate facts into a coherent picture. And what about algorithmic drift? That’s the problem where a model’s performance gets worse over time as the real-world patient population changes from what the model was trained on, a massive long-term risk in a dynamic environment like an ER.
The Gauntlet: Clinical Validation and Regulatory Hurdles
For VCs, the takeaway here is simple: if you’re going to invest in a clinical-facing LLM for a high-stakes job like triage, you need to demand ironclad clinical validation data. Getting these tools to market isn’t about a clever algorithm. It’s a long, expensive slog through regulatory frameworks built to protect patients. Unlike an admin tool, a triage application is usually classified as Software as a Medical Device (SaMD) by the FDA. That classification means a company has to go through a stringent premarket review, typically the 510(k) clearance pathway or, if it’s a totally new kind of device, the De Novo classification. A huge piece of getting that clearance is proving your device works and is safe with strong clinical trial data. If you check the FDA’s Center for Devices and Radiological Health (CDRH) database, you’ll see a steady increase in AI tools getting 510(k) clearances, with over 1,500 AI-enabled medical devices holding FDA marketing authorization expected by early 2026. A few are even for triage. In January 2026, for instance, Aidoc got FDA clearance for its CARE foundation model, which it calls a complete AI triage solution for spotting acute findings. A month earlier, in December 2025, a2z Radiology AI received clearance for its Unified-Triage system that flags urgent conditions on abdomen-pelvis CTs. But even with these wins, radiology still makes up the lion’s share of AI clearances (about 76%), and most other AI tools for the ER are for image analysis (like spotting a stroke on a CT scan) or are clinical decision support (CDS) tools that help a human make a better decision, not replace them. That distinction between CDS (which gives advice) and diagnostic AI (which makes a call) is everything when it comes to regulation and liability. Any company trying to sell an LLM for triage has to prove it follows Good Machine Learning Practice (GMLP) and has a solid Predetermined Change Control Plan (PCCP) to manage model updates without going through the entire premarket review from scratch each time, a necessity for LLMs. Without hitting these regulatory milestones and proving real clinical effectiveness, the chances of getting paid for these products are basically zero.
Methodology and Sources
This analysis is based on a review of peer-reviewed clinical studies on LLM performance from journals like JAMA Network Open. We also used public FDA regulatory guidance on clinical decision support software and AI/ML-enabled devices, along with data from the FDA CDRH 510(k) clearance database. Some perspective also comes from the official positions of groups like the American Medical Association. And while we mention companies like Epic Systems and Microsoft, the point here isn’t to evaluate their specific products but to look at the general, evidence-based challenges facing any LLM that’s put into a role like emergency triage. The goal is an objective look at where these models are failing right now.
Frequently Asked Questions
Why do general-purpose LLMs currently fail clinical validity in high-stakes applications like ED triage?
General-purpose LLMs excel at pattern matching on textual data but lack the inferential reasoning, causality understanding, and ability to manage diagnostic uncertainty required for nuanced clinical judgment. They rely on statistical relationships from vast datasets, which does not inherently confer clinical expertise or an understanding of physiological underpinnings. This leads to degradation in accuracy, especially for less common but critical conditions.
What evidence exists regarding the limitations of LLMs in clinical triage?
Peer-reviewed research, such as studies in JAMA Network Open, shows that LLMs exhibit concerning degradation in accuracy when evaluated against human nurses in simulated or retrospective triage scenarios. They can accurately classify common, low-acuity presentations, but their sensitivity and specificity significantly drop for less common but critical conditions, potentially leading to undertriage and tangible safety risks.
What regulatory considerations are critical for LLM-based clinical triage applications?
Clinical triage applications are often classified as Software as a Medical Device (SaMD) by regulatory bodies like the FDA. This classification necessitates stringent requirements for premarket review, typically through the 510(k) clearance pathway or De Novo classification. A critical component for these pathways is the demonstration of substantial equivalence or reasonable assurance of safety and effectiveness through robust clinical trial data.
How do LLMs differ in their effectiveness for administrative tasks versus direct clinical decision-making?
LLMs are powerful tools for administrative tasks like medical record summarization and prior authorization processing due to their ability to process and synthesize vast amounts of unstructured data. However, their success in these support roles should not be conflated with readiness for direct clinical decision-making, which demands nuanced clinical judgment, pattern recognition based on experience, and understanding of contextual patient factors that LLMs currently lack.
