AI trial-screening tools mostly ignore race. One label skewed results: homelessness.
A study running 5.3 million evaluations through nine large language models found eligibility judgments were largely stable across identity labels — except when a patient vignette mentioned homelessness.
As large language models get tested for tasks like screening patients for clinical trial eligibility, one open question is whether an AI's judgment about a patient shifts depending on details that shouldn't matter — a patient's race, their housing situation, their income. A study published in the Journal of the American Medical Informatics Association this month ran the experiment directly, and largely, though not entirely, found reassuring results [s1].
How the study was designed
Researchers built physician-validated clinical vignettes based on real Phase II and III US adult randomized controlled trial protocols from 2023 to 2024, then created 33 sociodemographic identity variants of each vignette — versions that differed only in the demographic labels attached to an otherwise identical patient case [s1]. Nine different large language models evaluated eligibility and related domains for each version [s1]. In total, the study covered 58 trial protocols and produced 5.3 million individual evaluations [s1].
What stayed stable
Across that large dataset, eligibility judgments were largely stable regardless of which sociodemographic identity was attached to a given vignette [s1]. Race and ethnicity specifically showed minimal effect on eligibility determinations once socioeconomic status was accounted for [s1] — a finding that, on its face, suggests the models were not simply pattern-matching demographic labels to eligibility outcomes in the way earlier generations of algorithmic bias research have documented in other clinical AI contexts.
Where it didn't
One label broke that pattern clearly: homelessness. Vignettes identifying a patient as homeless produced the largest negative shift in eligibility determinations of any identity variant tested, along with pronounced effects on how the models assessed the patient's likely treatment adherence, available resources, and trustworthiness [s1]. The researchers describe eligibility judgments as generally stable "except" for this case — making homelessness the clear outlier in an otherwise largely consistent dataset [s1].
Why the pattern matters
The study's authors draw a distinction between two different kinds of judgment a trial-screening AI has to make. On domains governed by explicit, written eligibility criteria — a specific lab value cutoff, a diagnosis code, an age range — the models applied the rules consistently regardless of the sociodemographic label attached [s1]. Disparities emerged specifically in domains that required the model to make an inference about a patient's likely behavior or resources rather than check a stated criterion [s1] — exactly the kind of soft judgment where a model trained on text reflecting real-world bias against homeless patients would be expected to reproduce that bias.
What this means for trial access
Clinical trial eligibility screening is increasingly a candidate task for AI assistance, given the volume of protocols and patient records involved. This study's finding — that explicit criteria were applied evenhandedly while judgment calls about a patient's behavior or resources were not — points to a specific, addressable failure mode rather than a blanket indictment of using LLMs for this purpose. The researchers frame their own findings as underscoring the need for careful deployment specifically to protect fair trial access for vulnerable patients, homeless patients in particular, rather than arguing against AI-assisted screening altogether [s1].
What the study doesn't establish
The vignettes were derived from real protocols but were not real patient encounters, and the study measured what the models said about hypothetical cases rather than tracking real-world trial enrollment outcomes. It also does not identify which specific text patterns in model training led to the homelessness effect, so it documents the disparity without fully explaining its mechanism. And a finding that race and ethnicity showed "minimal" effects after adjusting for socioeconomic status is not the same as showing zero effect before that adjustment — the raw, unadjusted picture across all 33 identity variants is not detailed in the portion of the study available here [s1].
Sources
- Sociodemographic bias in large language model clinical trial screening — Journal of the American Medical Informatics Association, 2026-05-12
Sources
- Sociodemographic bias in large language model clinical trial screening — Journal of the American Medical Informatics Association , May 12, 2026
Clinical AI has 63 ways to measure fairness and one built for clinical use
A Lancet Digital Health scoping review found the field's fairness metrics fragmented and rarely clinically validated. A second review found that most studies don't measure fairness at all.
A computer that grades each colonoscopy raised how often endoscopists found adenomas
A Danish stepped-wedge trial gave endoscopists automated feedback on their technique after every procedure. Adenoma detection rose from 43.4% to 48.6% — a different tool from real-time polyp AI.
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.