Listen to this article · 8 min listen

The promise of artificial intelligence in healthcare is often framed through lenses of efficiency and precision, yet a critical, often overlooked, dimension is its potential to exacerbate existing health inequities. When AI systems are designed using flawed assumptions, particularly those rooted in economic proxies for health, they risk embedding and amplifying racial bias, leading to documented failures in patient care. This phenomenon, which we term the “proxy variable problem,” poses a significant threat to patient safety and demands rigorous scrutiny from all stakeholders, from clinical informaticists to regulatory officers.

The Proxy Variable Problem: When Cost Becomes a Stand-in for Need

At the heart of algorithmic racial bias in healthcare AI lies the dangerous assumption that readily available economic data, such as past healthcare spending, can serve as an accurate proxy for an individual’s true health needs. This approach, while seemingly logical from a cost-efficiency perspective for entities like large health systems or insurers, fundamentally misunderstands the systemic barriers that prevent certain populations from accessing adequate care. Consequently, algorithms trained on such data learn to associate lower past spending with lower health risk, inadvertently penalizing individuals who have historically faced socioeconomic or racial discrimination in healthcare access. A seminal investigation into this issue, led by Ziad Obermeyer, a distinguished associate professor at UC Berkeley, and his colleagues, uncovered a stark example within an algorithm widely used by Optum, a subsidiary of UnitedHealth Group. This algorithm was designed to predict which patients would benefit most from “high-risk care management” programs. Instead of directly predicting illness severity or future health complications, the algorithm used healthcare costs as its primary outcome variable. The study, published in Science in 2019, revealed that the algorithm systematically assigned lower risk scores to Black patients than to white patients, even when both groups had the same level of chronic illness Study on racial bias in healthcare algorithms. This bias meant that Black patients, despite being sicker, were less likely to be flagged for proactive care interventions. The researchers estimated that, had the algorithm been unbiased, it would have more than doubled the number of Black patients identified for these crucial programs. The core of the problem, as highlighted by Obermeyer’s team, was that Black patients, due to systemic healthcare disparities, typically incurred lower healthcare costs for a given level of illness compared to white patients. Factors such as less frequent doctor visits, delayed diagnoses, and unequal access to specialized care contribute to this disparity. By optimizing for cost, the algorithm effectively encoded and perpetuated these existing racial biases, leading to an undertriage of Black patients with significant health needs.

Unpacking the Mechanisms of Algorithmic Bias

The case of Optum’s algorithm illustrates how seemingly objective metrics can become conduits for deeply entrenched societal biases. Ruha Benjamin, a professor at Princeton University and a leading voice on race, technology, and justice, eloquently describes this as the “New Jim Code”, where technological systems, often cloaked in claims of neutrality, can reproduce and even amplify racial inequalities. Her work, and the findings from institutions like UC Berkeley and Harvard T.H. Chan School, underscores that the problem is not merely about flawed data, but about the flawed assumptions embedded in the design and deployment of these AI systems. The relationship between cost as a proxy for health needs and algorithmic racial bias is not incidental; it is, as research suggests, the most widespread mechanism for such bias in healthcare AI. When developers prioritize easily quantifiable metrics like past spending, they often overlook the complex social determinants of health and the historical context of healthcare access. This leads to a situation where AI, rather than improving equity, becomes a tool for perpetuating disparities. The economic implications are also significant. For public health, the undertreatment of specific populations can lead to worsening health outcomes, increased emergency care utilization, and higher long-term costs Economic impact of health disparities. For health systems, relying on biased algorithms can undermine their mission to provide equitable care and expose them to significant reputational and regulatory risks.

Regulatory Imperatives and the Path to Responsible AI

The documented failures arising from the proxy variable problem underscore the urgent need for robust regulatory frameworks and proactive fairness auditing. The FDA’s Software as a Medical Device (SaMD) Framework, while crucial for ensuring the safety and efficacy of AI-driven medical devices, must evolve to explicitly address algorithmic fairness as a core component of validation. Similarly, the FTC’s Algorithmic Fairness initiatives provide a foundational understanding of the consumer protection aspects of AI, which are directly applicable to patient safety in healthcare. Institutions like UC Berkeley, Princeton University, and Harvard T.H. Chan School are at the forefront of researching and advocating for responsible AI development. Their work emphasizes that clinical validation for AI in healthcare must extend beyond traditional performance metrics (like accuracy or sensitivity) to include rigorous assessments of fairness across different demographic groups. This requires moving beyond simple aggregate performance to understand how an algorithm performs for specific subgroups, particularly those historically marginalized. For patient safety advocates, FDA and regulatory officers, and clinical informaticists, the implications are clear. The integration of fairness auditing into existing regulatory frameworks is not merely a desirable add-on but a critical necessity. This involves pre-market assessments that scrutinize the training data for representativeness and bias, as well as post-market surveillance mechanisms that continuously monitor for algorithmic drift and differential performance across patient populations. Such auditing should draw insights from organizations like the Commonwealth Fund or RAND, which have extensively studied health equity and quality improvement. Evidence formats critical for demonstrating safety and efficacy must encompass not only clinical outcomes but also equity outcomes, ensuring that AI benefits all patients, not just some.

Toward Algorithmic Equity and Patient Safety

The pervasive issue of cost-based algorithms encoding racial bias represents a profound challenge to the promise of AI in healthcare. The documented failures, exemplified by the Optum/UnitedHealth Group case and highlighted by researchers like Ziad Obermeyer and Ruha Benjamin, serve as a stark reminder that unchecked AI can amplify existing societal inequities. The solution lies not in abandoning AI, but in a concerted effort to build and deploy AI systems responsibly, with an unwavering commitment to fairness and equity. This demands that we critically examine the underlying assumptions in algorithm design, prioritize diverse and representative data, and establish robust regulatory oversight that explicitly mandates fairness auditing. Only then can we ensure that AI truly serves as a tool for improving health outcomes for all, rather than inadvertently perpetuating harm.

Frequently Asked Questions

A5: How does the ‘proxy variable problem’ endanger patient safety?

The ‘proxy variable problem’ endangers patient safety by using economic data, like past healthcare spending, as a stand-in for true health needs. This can lead to algorithms associating lower past spending with lower health risk, inadvertently penalizing individuals who have faced systemic barriers to care. Consequently, sicker patients from certain demographic groups may be less likely to be identified for crucial proactive care interventions, leading to documented failures in patient care.

A3: What regulatory gaps or considerations does the ‘proxy variable problem’ highlight for AI in healthcare?

The ‘proxy variable problem’ highlights the urgent need for robust regulatory frameworks and proactive fairness auditing in healthcare AI. While frameworks like the FDA’s Software as a Medical Device (SaMD) are important, they must evolve to explicitly address algorithmic fairness as a core component of validation. Regulatory bodies need to ensure that clinical validation for AI extends beyond traditional performance metrics to include rigorous assessments of fairness across different demographic groups.

A2: How can clinical informaticists mitigate the ‘proxy variable problem’ when designing or implementing AI systems?

Clinical informaticists can mitigate the ‘proxy variable problem’ by avoiding the dangerous assumption that readily available economic data, such as past healthcare spending, accurately represents an individual’s true health needs. They should move beyond prioritizing easily quantifiable metrics and instead consider the complex social determinants of health and the historical context of healthcare access. This involves designing AI systems that do not embed flawed assumptions rooted in economic proxies for health, which can amplify racial bias.

A5: Can you provide a specific example of how the ‘proxy variable problem’ has led to biased patient care?

A seminal investigation by Ziad Obermeyer and colleagues found that an algorithm used by Optum, designed to predict patients for ‘high-risk care management’ programs, used healthcare costs as its primary outcome variable. This algorithm systematically assigned lower risk scores to Black patients than to white patients, even when both groups had the same level of chronic illness. As a result, Black patients, despite being sicker, were less likely to be flagged for proactive care interventions, leading to an undertriage of their significant health needs.