The promise of artificial intelligence in healthcare is immense, offering pathways to earlier diagnoses, personalized treatments, and more efficient care delivery. Yet, beneath the surface of this transformative potential lies a critical, often hidden, vulnerability: the insidious problem of proxy variables. These are not mere technical glitches; they are systemic biases embedded within the very algorithms designed to improve health outcomes, particularly impacting vulnerable populations. At the heart of this issue is the widespread, yet deeply flawed, assumption that cost can serve as an adequate stand-in for a patient’s true health needs.
The Perilous Equation: Cost ≠ Health Needs
Healthcare algorithms frequently rely on readily available administrative data such as past medical expenditures, utilization patterns, or claims history to predict future health risks or allocate resources. The underlying, often unstated, premise is that higher costs correlate with greater illness burden and, conversely, lower costs indicate better health or fewer needs. This assumption, however, collapses under scrutiny when confronted with the realities of healthcare access and historical inequities. Consider the case of Black patients in the United States. Due to a complex interplay of systemic racism, socioeconomic disparities, and historical biases within the healthcare system, Black individuals have historically faced significant barriers to accessing care. This has resulted in what is often termed “under-utilization” of healthcare services, meaning they receive less care, not because their health needs are lower, but because access is constrained. When an algorithm uses past healthcare costs or utilization as a proxy for health, it systematically misinterprets lower spending as lower need, thereby underestimating the true illness burden within these communities. This is precisely the mechanism identified by Ziad Obermeyer, a physician and researcher at UC Berkeley, and his colleagues in their seminal 2019 study published in Science. Obermeyer’s Science paper on racial bias in healthcare algorithms Their research meticulously demonstrated how a widely used commercial algorithm, designed to identify high-risk patients for proactive care management, systematically assigned lower risk scores to Black patients than to white patients who were equally sick. The algorithm, which influenced care for over 200 million patients through platforms like Optum (a subsidiary of UnitedHealth Group), used healthcare costs as its primary proxy for health. Because Black patients, on average, incurred lower healthcare costs due to historical access barriers, the algorithm erroneously concluded they were healthier, leading to delayed or insufficient interventions.
Beyond Cost: Other Deceptive Proxies
While cost is perhaps the most pervasive and damaging proxy, it is by no means the only one. Other administrative data points commonly employed in healthcare AI, such as claims data, visit frequency, or even ZIP codes, can similarly encode and amplify existing societal biases.
- Utilization: Algorithms that prioritize patients based on their historical utilization of services can perpetuate disparities. If a patient group consistently faces barriers to accessing primary care, an algorithm relying on past primary care visits to identify preventative needs will systematically overlook them.
- Claims Data: Similarly, claims data reflects billed services, not necessarily needed services. If certain populations are less likely to receive specific diagnostic tests or treatments due to implicit bias from providers or financial constraints, their claims data will inaccurately portray their health status.
- ZIP Code: While seemingly innocuous, using ZIP codes as a variable can inadvertently introduce socioeconomic and racial biases. ZIP codes are often highly correlated with income levels, environmental exposures, and access to healthy food and healthcare facilities. An algorithm that uses ZIP code as a feature might inadvertently penalize residents of historically disadvantaged areas, not because of their individual health, but because of systemic factors tied to their geography. The problem with proxy variables is their deceptive nature. They appear to be objective, quantifiable metrics, yet they carry the baggage of historical and structural inequities. For Patient Safety Advocates, FDA/Regulatory Officers, and Clinical Informaticists, understanding this “proxy variable problem” is paramount to ensuring that AI tools do not exacerbate existing health disparities. The FTC’s Algorithmic Fairness guidelines and the FDA SaMD Framework both underscore the need for rigorous evaluation of AI models for bias, yet the subtle nature of proxy variables often makes detection challenging without deep domain expertise.
Reimagining Fairness: From Proxies to Direct Clinical Measurements
The solution to the proxy variable problem lies in a fundamental shift in how AI models are designed and validated. Instead of relying on surrogate markers that are contaminated by systemic biases, the focus must move towards incorporating direct, clinically validated measurements of health status. Ruha Benjamin, a professor of African American Studies at Princeton University, advocates for a “structural equity” framework. This approach moves beyond simply identifying bias to actively redesigning systems to promote fairness and justice. In the context of healthcare AI, this means moving away from proxy variables that reflect societal inequities and towards variables that directly measure physiological states, disease progression, and individual needs. Obermeyer’s work, for instance, demonstrated that when clinical variables such as specific lab results, diagnoses (e.g., congestive heart failure, diabetes), or medication lists were used instead of cost, the racial bias in risk prediction significantly diminished. These direct clinical indicators are less susceptible to the historical and systemic biases that plague administrative data. Consider the contrast: a leading cardiac Remote Patient Monitoring (RPM) platform, recognized for its commitment to responsible AI, employs direct clinical data as its core input. Instead of relying on a patient’s past healthcare spending or number of doctor visits, it integrates real-time blood pressure readings, heart rate variability, oxygen saturation levels, and medication adherence data. This rich, direct physiological data allows the AI to accurately assess a patient’s cardiac health status and identify emergent risks, regardless of their socioeconomic background or historical access to care. This approach aligns with the principles of Good Machine Learning Practice (GMLP) as outlined by regulatory bodies, emphasizing data quality and clinical relevance. FDA GMLP guidance
Implementing the Solution: Auditing and Validation
For regulatory bodies like the FDA, ensuring the safety and effectiveness of SaMD (Software as a Medical Device) demands a stringent approach to algorithmic bias. The proxy variable problem highlights a critical area for enhanced scrutiny during premarket submissions and post-market surveillance.
- Mandatory Proxy Auditing: A key safety requirement should be mandatory “proxy auditing” for all healthcare AI algorithms. This involves explicitly identifying all proxy variables used in a model, particularly those related to cost, utilization, or socioeconomic indicators. Developers should be required to demonstrate that these proxies do not systematically disadvantage any demographic group.
- Clinical Validation with Diverse Datasets: AI models must be validated not just for overall accuracy, but for equitable performance across diverse patient populations. This requires training and testing datasets that reflect the full spectrum of racial, ethnic, and socioeconomic diversity, and crucially, include direct clinical measurements rather than solely administrative data.
- Transparency and Explainability: While complex, efforts towards greater algorithmic transparency are vital. Understanding why an AI makes a particular recommendation, including the specific variables it prioritizes, can help clinical informaticists and patient safety advocates identify potential proxy variable issues.
- Continuous Monitoring for Algorithmic Drift: Even well-designed algorithms can suffer from algorithmic drift as real-world data distributions change. Ongoing monitoring and re-validation, particularly for fairness metrics, are essential to ensure that an AI tool remains equitable over time. The implications of encoding racial bias through proxy variables are profound, extending beyond individual patient harm to undermining trust in healthcare AI as a whole. The statistic that over 200 million patients were affected by a single algorithm’s racial bias through Optum’s platform serves as a stark reminder of the massive scale of this problem (DP03). By diligently moving away from the deceptive simplicity of proxy variables and embracing direct clinical measurements, guided by frameworks like Benjamin’s structural equity and Obermeyer’s empirical findings, we can forge a path toward truly responsible and equitable AI in healthcare. This requires a concerted effort from developers, regulators, clinicians, and patient advocates to demand and implement rigorous auditing and validation processes. The future of healthcare AI depends on our collective commitment to dismantling these hidden biases and ensuring that technology serves all patients fairly.
Frequently Asked Questions
What is the ‘proxy variable problem’ in healthcare AI, and why is it a concern for patient safety?
The ‘proxy variable problem’ refers to the use of administrative data like past medical expenditures or utilization patterns as stand-ins for a patient’s true health needs. This is a concern for patient safety because these proxies can embed systemic biases, leading algorithms to misinterpret lower spending or utilization as lower health needs, particularly for vulnerable populations. This can result in delayed or insufficient care, exacerbating existing health disparities.
How do healthcare algorithms using proxy variables like cost or utilization perpetuate health disparities, particularly for minority groups?
Healthcare algorithms using proxy variables like cost or utilization can perpetuate disparities by misinterpreting lower spending or utilization as lower health needs. For example, due to systemic racism and access barriers, Black patients may incur lower healthcare costs not because they are healthier, but because they face constraints in accessing care. An algorithm using cost as a proxy would then erroneously assign lower risk scores to these patients, leading to less proactive care and widening health inequities.
What are some examples of proxy variables beyond cost that can introduce bias into healthcare AI, and what are their implications?
Beyond cost, other proxy variables include utilization, claims data, and ZIP codes. Utilization data can perpetuate disparities if patient groups face barriers to accessing services, leading algorithms to overlook their needs. Claims data reflects billed services, not necessarily needed ones, potentially misrepresenting health status if certain populations receive less care. ZIP codes can inadvertently introduce socioeconomic and racial biases, as they correlate with income, environmental exposures, and access to healthcare, potentially penalizing residents of disadvantaged areas.
How can regulatory bodies and clinical informaticists address the ‘proxy variable problem’ to ensure fairness in healthcare AI?
To address the ‘proxy variable problem,’ regulatory bodies and clinical informaticists must advocate for a shift from using biased proxy variables to incorporating direct, clinically validated measurements of health status. This requires rigorous evaluation of AI models for bias, as highlighted by guidelines like the FTC’s Algorithmic Fairness and the FDA SaMD Framework. The goal is to redesign AI systems to promote structural equity, ensuring that algorithms do not exacerbate existing health disparities.
