ANALYSIS

Clinical AI has 63 ways to measure fairness and one built for clinical use

A Lancet Digital Health scoping review found the field's fairness metrics fragmented and rarely clinically validated. A second review found that most studies don't measure fairness at all.

"Is this algorithm fair?" sounds like a question with an answer. Two reviews published this year establish that in clinical AI it is currently a question with sixty-three answers, most of them untested, and that in practice almost nobody asks it.

The metrics problem

A scoping review published in The Lancet Digital Health on 28 May set out to identify and critically appraise the fairness metrics used in clinical predictive AI models [s1]. The authors defined a fairness metric as one quantifying whether a model discriminates societally against individuals or groups defined by sensitive attributes [s1].

They searched five databases for literature published between 2014 and 2024, screened 820 records, included 42 studies and extracted 63 fairness metrics [s1]. The search was limited to studies published in English [s1].

Sixty-three metrics from 42 studies is close to one and a half distinct definitions of fairness per paper. The review classified them by performance dependency, model output level and base performance metric, and describes the resulting landscape as fragmented, with inadequate clinical validation and over-reliance on threshold-dependent metrics [s1].

The sharpest number in the review is this: of the 63 metrics, 19 were explicitly developed for health care, and among those 19, only one was developed for clinical use [s1].

That distinction — developed for health care versus developed for clinical use — is doing real work. A metric can be built on medical data and still be a research instrument, computed after the fact on a validation set, with no defined action attached to any particular value. One metric out of 63 was built to be used at the point where a decision gets made.

The review also identifies gaps in uncertainty quantification, intersectionality, and real-world applicability, and concludes that future work on clinical predictive AI models should prioritise clinically meaningful metrics [s1].

The measurement problem

The second review, published in npj Digital Medicine on 14 July, asks a blunter question: how often is fairness measured at all [s2].

It examined fairness in multimodal AI-based clinical decision support systems — models combining more than one data type — with rule-based and knowledge-based systems out of scope [s2]. The review was registered with PROSPERO [s2].

From 3,059 search results, 160 articles included fairness evaluations, of which 11% (18 of 160) used multimodal data, and across those, 29 different fairness measurement techniques appeared [s2].

The second search is the damaging one. From 1,422 results relating to two focus areas — chest X-rays and sepsis — the review identified 88 studies using multimodal data for clinical decision support [s2]. Fairness evaluations appeared in 8% of them: 7 of 88 [s2].

And 83% of those studies — 74 of 88 — included data on sensitive attributes that would have enabled a fairness evaluation [s2].

Reading the two together

The combination is what makes this a story rather than two literature reviews. Nine out of ten studies in two of the most-modelled clinical domains did not evaluate fairness despite holding the data needed to do so [s2]. Among the minority that did, the field offers 63 metrics, fragmented across three classification axes, with a single one designed for use in clinical practice [s1].

So the problem is not primarily that researchers are choosing bad fairness metrics. It is that the default is no metric, and that a researcher who does want to check has to select from a menu with no clinical consensus behind it and little validation attached to any option.

Both papers converge on the same recommendation from different directions. The Lancet Digital Health authors call for prioritising clinically meaningful metrics [s1]; the npj authors conclude that research on multimodal clinical decision support would benefit from guidance and tools supporting fairness evaluations [s2].

What these reviews are

Both are reviews of the published literature, which is a specific and limited window. They describe what researchers report in papers, not what deployed systems do inside health systems. A model validated internally at a hospital, with a fairness audit that never gets published, is invisible to both. The direction of that bias is not obvious: unpublished audits could be more common than the literature suggests, or considerably less.

The Lancet Digital Health review's English-language restriction [s1] is a further constraint on generalising its metric census to the global field.

Why the timing matters

Fairness metrics are the instrument that would detect a specific failure mode — a model that performs acceptably in aggregate while performing worse for a subgroup. That failure is invisible to overall accuracy figures, which is precisely why it needs its own measurement.

As clinical prediction tools move into routine use, the question of who validates them for subgroup performance, and against what standard, becomes an operational one rather than an academic one. These two reviews establish that the standard does not currently exist in usable form, and that the measurement is usually skipped.

What to watch

Whether any regulator, guideline body or journal specifies which fairness metric a clinical prediction model should report. Whether reporting standards for prediction models add a subgroup performance requirement. And whether the one clinical-use metric the Lancet Digital Health review identified [s1] gets taken up, or stays at one.

Sources

  1. [s1] Critical appraisal of fairness metrics for artificial intelligence-based clinical prediction models: a scoping review. The Lancet Digital Health, published online 28 May 2026. https://doi.org/10.1016/j.landig.2026.101001
  2. [s2] Fairness in multimodal machine learning applications in clinical decision support: a systematic review. npj Digital Medicine, 14 July 2026. PROSPERO CRD42024579923. https://doi.org/10.1038/s41746-026-03000-x

Sources

  1. Critical appraisal of fairness metrics for artificial intelligence-based clinical prediction models: a scoping reviewThe Lancet Digital Health , May 28, 2026
  2. Fairness in multimodal machine learning applications in clinical decision support: a systematic reviewnpj Digital Medicine , July 14, 2026

More on

Related coverage