ANALYSIS

AI trial-screening tools mostly ignore race. One label skewed results: homelessness.

A study running 5.3 million evaluations through nine large language models found eligibility judgments were largely stable across identity labels — except when a patient vignette mentioned homelessness.

As large language models get tested for tasks like screening patients for clinical trial eligibility, one open question is whether an AI's judgment about a patient shifts depending on details that shouldn't matter — a patient's race, their housing situation, their income. A study published in the Journal of the American Medical Informatics Association this month ran the experiment directly, and largely, though not entirely, found reassuring results [s1].

How the study was designed

Researchers built physician-validated clinical vignettes based on real Phase II and III US adult randomized controlled trial protocols from 2023 to 2024, then created 33 sociodemographic identity variants of each vignette — versions that differed only in the demographic labels attached to an otherwise identical patient case [s1]. Nine different large language models evaluated eligibility and related domains for each version [s1]. In total, the study covered 58 trial protocols and produced 5.3 million individual evaluations [s1].

What stayed stable

Across that large dataset, eligibility judgments were largely stable regardless of which sociodemographic identity was attached to a given vignette [s1]. Race and ethnicity specifically showed minimal effect on eligibility determinations once socioeconomic status was accounted for [s1] — a finding that, on its face, suggests the models were not simply pattern-matching demographic labels to eligibility outcomes in the way earlier generations of algorithmic bias research have documented in other clinical AI contexts.

Where it didn't

One label broke that pattern clearly: homelessness. Vignettes identifying a patient as homeless produced the largest negative shift in eligibility determinations of any identity variant tested, along with pronounced effects on how the models assessed the patient's likely treatment adherence, available resources, and trustworthiness [s1]. The researchers describe eligibility judgments as generally stable "except" for this case — making homelessness the clear outlier in an otherwise largely consistent dataset [s1].

Why the pattern matters

The study's authors draw a distinction between two different kinds of judgment a trial-screening AI has to make. On domains governed by explicit, written eligibility criteria — a specific lab value cutoff, a diagnosis code, an age range — the models applied the rules consistently regardless of the sociodemographic label attached [s1]. Disparities emerged specifically in domains that required the model to make an inference about a patient's likely behavior or resources rather than check a stated criterion [s1] — exactly the kind of soft judgment where a model trained on text reflecting real-world bias against homeless patients would be expected to reproduce that bias.

What this means for trial access

Clinical trial eligibility screening is increasingly a candidate task for AI assistance, given the volume of protocols and patient records involved. This study's finding — that explicit criteria were applied evenhandedly while judgment calls about a patient's behavior or resources were not — points to a specific, addressable failure mode rather than a blanket indictment of using LLMs for this purpose. The researchers frame their own findings as underscoring the need for careful deployment specifically to protect fair trial access for vulnerable patients, homeless patients in particular, rather than arguing against AI-assisted screening altogether [s1].

What the study doesn't establish

The vignettes were derived from real protocols but were not real patient encounters, and the study measured what the models said about hypothetical cases rather than tracking real-world trial enrollment outcomes. It also does not identify which specific text patterns in model training led to the homelessness effect, so it documents the disparity without fully explaining its mechanism. And a finding that race and ethnicity showed "minimal" effects after adjusting for socioeconomic status is not the same as showing zero effect before that adjustment — the raw, unadjusted picture across all 33 identity variants is not detailed in the portion of the study available here [s1].

Sources

Sources

  1. Sociodemographic bias in large language model clinical trial screeningJournal of the American Medical Informatics Association , May 12, 2026
Related coverage