WHAT THE STUDY ACTUALLY SAYS

Skin-lesion AI is accurate on paper but weaker on darker skin, two 2026 reviews find

Pooled diagnostic accuracy is high in specialist settings, but performance falls on Fitzpatrick IV–VI skin and on smartphone images, and most studies never report skin tone at all.

Skin-lesion AI accuracy (AUROC) by Fitzpatrick skin typeFitzpatrick I–III (lighter skin): 0.89; Fitzpatrick IV–VI (darker skin): 0.8200.450.9Fitzpatrick I–III (lighter skin)0.89Fitzpatrick IV–VI (darker skin)0.82
Skin-lesion AI accuracy (AUROC) by Fitzpatrick skin type
GroupValue (value)
Fitzpatrick I–III (lighter skin)0.89
Fitzpatrick IV–VI (darker skin)0.82
Skin-lesion AI accuracy (AUROC) by Fitzpatrick skin type Pooled AUROC across studies reporting skin-tone-stratified results; higher is better. Source: Medicina

Artificial intelligence for reading skin lesions is accurate in the settings where it has mostly been tested, but two 2026 meta-analyses agree it performs worse on darker skin and on smartphone photographs — and that most published studies never record skin tone at all [s1][s2]. On the head-to-head question that matters most, the evidence does not show autonomous AI beating a clinician with a dermatoscope; the two perform comparably [s2].

What the pooled evidence shows

A systematic review and meta-analysis in Medicina, published in December 2025, drew on 18 studies covering more than 70,000 test images across specialist clinics, community or primary care, and smartphone settings [s1]. Pooled sensitivity was high at 0.91 (95% CI 0.74–0.97), but pooled specificity was much lower at 0.64 (95% CI 0.47–0.78), and the overall area under the ROC curve was 0.88 (95% CI 0.87–0.90) [s1].

The low specificity is the quieter of those numbers and the more consequential in a screening context: it means a system that catches most cancers also flags a large share of benign lesions, generating referrals and biopsies. High sensitivity with modest specificity is a triage profile, not a diagnostic one — useful for deciding who sees a dermatologist, not for ruling cancer in or out.

Where accuracy falls away

The review's central equity finding is a gap by skin tone. Pooled accuracy was AUROC 0.89 on Fitzpatrick I–III (lighter) skin against 0.82 on Fitzpatrick IV–VI (darker) skin, a difference of 0.07 that the authors report as statistically significant (p<0.01) [s1]. Performance also fell by setting: AUROC was 0.90 in specialist clinics, 0.85 in community care and 0.81 on smartphone images [s1] — the last being exactly the direct-to-consumer use case these tools are most often marketed for.

The reporting gap underneath those numbers is the deeper problem. Only 6 of the 18 studies provided any skin-tone-stratified outcomes [s1]. When most of the literature does not record the variable on which performance visibly differs, the true size of the disparity is unknown, and a headline accuracy figure averages over populations the model may serve unequally. This is the same failure mode documented in other consumer sensors — see our reporting on pulse oximeters and skin tone and on why fairness metrics for clinical AI are contested.

AI versus the clinician

A second independent meta-analysis, in Frontiers in Medicine in July 2026, asked whether AI actually outperforms dermoscopy for pigmented lesions [s2]. Across 10 studies and 17 diagnostic arms, dermoscopy by clinicians pooled to a sensitivity of 0.773 (95% CI 0.648–0.863) and specificity of 0.793 (95% CI 0.673–0.877); standalone AI pooled to sensitivity 0.757 (95% CI 0.428–0.928) and specificity 0.859 (95% CI 0.619–0.958) [s2]. The confidence intervals overlap heavily — the AI sensitivity interval runs from 0.43 to 0.93 — and the authors report no clear separation between the summary ROC curves [s2].

Their conclusion is deliberately deflationary: current evidence does not substantiate the consistent superiority of autonomous AI, which should be treated as a supplementary tool rather than a replacement for clinical assessment [s2]. The one arm testing AI-assisted clinicians reported a sensitivity of 1.000 and specificity of 0.837 — encouraging, but a single arm, from which nothing general can be concluded [s2].

The limits behind the numbers

Both reviews flag the same generalisability problem. The Frontiers authors judged most included studies at high risk of selection bias because they enrolled patients whose lesions were already suspicious rather than unselected screening populations [s2] — a design that inflates apparent accuracy relative to real first-line use. And both note that darker skin tones are underrepresented in the training datasets, which is the mechanism, not merely a correlate, of the performance gap [s1][s2].

Two further caveats: the Medicina pooled estimates rest on the six studies with complete data, and Fitzpatrick reporting was inconsistent throughout [s1]; the Frontiers evidence on AI-assisted assessment is a single arm [s2]. Neither review is a substitute for a large prospective trial in an unselected population, stratified by skin tone — which is what the field still lacks.

What it means

For a reader, the practical translation is narrow. These tools can help sort lesions for review, and in specialist hands on lighter skin they read well; they have not been shown to diagnose skin cancer as well as, let alone better than, an examining clinician, and they are least validated in precisely the darker-skinned and smartphone populations where they are most likely to be used unsupervised [s1][s2]. A benign-looking result from a phone app is not a clearance. For what changing or worrying lesions actually warrant attention, see our explainer on melanoma screening and warning signs and on Australia's readiness for risk-based melanoma screening; for a broader look at general-purpose diagnostic AI, see the DxDirector agentic-AI study.

Sources

Sources

  1. Equity and Generalizability of Artificial Intelligence for Skin-Lesion Diagnosis Using Clinical, Dermoscopic, and Smartphone Images: A Systematic Review and Meta-AnalysisMedicina , December 10, 2025
  2. Artificial intelligence vs. dermoscopy for malignancy risk stratification of pigmented skin lesions: a systematic review and meta-analysis for public healthFrontiers in Medicine , July 28, 2026

More on

Related coverage