Medical imaging foundation models leak enough signal to re-identify patients
Off-the-shelf features from two published models matched retinal scans to the right patient 78% to 86% of the time. Fine-tuned for the task, one model reached 99.5% on OCT.
De-identified medical images are shared on the assumption that stripping the metadata is enough. A study published this month tests that assumption against the current generation of imaging foundation models, and the answer is uncomfortable: the numerical feature vectors those models produce carry enough patient-specific signal to match one image to another from the same person, often on the first try [s1].
What was tested
The authors took two published, off-the-shelf foundation models — RETFound for ophthalmic images and CXR-Foundation for chest radiographs — froze them, and used them only as feature extractors [s2]. No fine-tuning, no privacy attack designed into the model; just the representations these models already produce for downstream tasks.
The procedure was simple. Treat each image as a query, compute feature similarity against every other image in the dataset, and check whether the most similar image belongs to the same patient. Images taken on the same day as the query were excluded from comparison, so the task was matching a person across separate encounters, not matching duplicates from one visit [s2].
Five datasets were used. In ophthalmology: 33,697 Topcon colour fundus photographs from 2,796 patients (CORIS-CFP) and 332,794 Spectralis OCT B-scans from 1,000 patients (CORIS-OCT), both from the Denver area, plus the public GRAPE dataset — 631 fundus photographs from 144 patients in Hangzhou, China [s2]. In radiology: 106,563 chest radiographs from 60,020 patients in the Boston area (MGH), and 106,473 selected images from 39,749 patients in the public MIDRC collection, drawn from the NIH ChestX-ray8 and Stanford CheXpert datasets [s2].
The result
From frozen features alone, first-choice re-identification at image level was 40.2% for CORIS-CFP, 46.3% for CORIS-OCT, 38.9% for GRAPE, 25.9% for MGH and 50.1% for MIDRC [s2]. At patient level — pooling a patient's images — the figures were 78.1%, 86.1%, 54.7%, 30.7% and 48.5% [s2]. Within the top ten candidates, patient-level rates reached 86.5% and 89.8% for the two internal ophthalmology datasets [s2].
(The paper's abstract gives 40.3% for CORIS-CFP where its results table gives 40.2% [s1][s2]. The difference is immaterial; the table figure is the one to quote.)
The single strongest determinant of success was how many separate encounters a patient had. Restricting the radiology datasets by minimum number of time points, image-level first-choice re-identification rose from 18.2% at two time points to 53.8% at five or more for MGH, and from 18.6% at two to 70.3% at seven or more for MIDRC [s2]. At seven or more time points, MIDRC patient-level first-choice re-identification reached 96.6%, and 99.4% within the top ten [s2].
That is the practical shape of the risk. A patient with one scan in a shared dataset is fairly safe. A patient with a longitudinal series — which is exactly what makes a dataset valuable for research — is not.
What happens if someone tries
The frozen-feature results describe an incidental property. The authors then fine-tuned models explicitly for re-identification, to establish an upper bound [s2].
Fine-tuning raised image-level first-choice rates to 82.3% for CORIS-CFP, 94.0% for CORIS-OCT, 63.7% for MGH and 74.0% for MIDRC [s2]. At patient level, OCT reached 99.5% first-choice, with average precision of 100 [s2].
The gap between frozen and fine-tuned performance is instructive. At patient level the frozen features were already close to the trained ceiling — 81.8% against 95.0% for fundus photographs, 87.0% against 99.5% for OCT — while at image level the gap stayed wide [s2]. In other words, the identity signal is largely present in the general-purpose representation; training mostly helps when only a single image is available.
Is it just demographics?
One hypothesis is that these models are simply encoding age, sex and race, and that re-identification is a byproduct of demographic clustering. The authors tested it by training linear probes on the frozen features and then splitting performance by whether an image had been correctly re-identified [s2].
The probes worked well on their own: gender AUC-ROC of 76.9% on CORIS-CFP, 69.3% on CORIS-OCT and 95.4% on MGH; race prediction above 95% AUC-ROC; and age prediction with R² around 0.7 [s2].
Stratified by re-identification status, the ophthalmology datasets showed the expected pattern — better demographic prediction on correctly re-identified images, with gender AUC-ROC of 82.1% versus 76.8% on CORIS-CFP, for example — but the authors describe the differences as modest, and conclude that more information contributes to re-identification than the four demographic features analysed [s2]. On MIDRC the pattern did not appear at all, which they suggest means radiological re-identification does not rely on those features [s2].
So the leak is not reducible to demographics. Something more individual — anatomy, device signature, or both — is being encoded.
What this does and does not mean
It does not mean an attacker can put a name to a scan. Re-identification here means linking two images to the same person within a dataset. Turning that into an identity requires an external reference — a second dataset where the person is named, or an image the attacker already possesses.
That is not a reassuring caveat so much as a description of the standard re-identification threat model. Linkage is the mechanism by which de-identified data becomes identified, and this study establishes that the linkage step is cheap: it requires no special model, only the published one, and a nearest-neighbour search.
The limits worth naming are in the datasets. Both internal cohorts and MIDRC are homogeneous in race and ethnicity, with most patients Caucasian and non-Hispanic, and each reflects the population its institution serves [s2]. Performance on GRAPE was lower than on the internal fundus dataset, which the authors attribute to fewer time points per patient and only 100 patients with longitudinal data against nearly 1,000 internally [s2]. And every result is retrospective, on curated research datasets rather than on whatever gets posted publicly.
Why it matters now
Foundation-model features are increasingly shared as a convenience — smaller than images, apparently abstract, seemingly safe to release for downstream model development or federated work. This paper's finding is that the feature vector is not an anonymising transformation. It is a compressed representation that retains identity.
The authors' own conclusion is that imaging features extracted from foundation models in ophthalmology and radiology include information that can lead to patient re-identification [s1]. Anyone treating embedding release as equivalent to de-identification now has a number to argue against.
Sources
- [s1] Re-identification of patients from imaging features extracted by foundation models — npj Digital Medicine, published online 22 July 2025. https://doi.org/10.1038/s41746-025-01801-0
- [s2] Same paper, open-access full text and tables — Europe PMC, PMC12283961. https://europepmc.org/article/PMC/PMC12283961
Sources
- Re-identification of patients from imaging features extracted by foundation models — npj Digital Medicine , July 22, 2025
- Re-identification of patients from imaging features extracted by foundation models — open-access full text and tables — Europe PMC (PMC12283961) , July 22, 2025
An AI eye copilot raised correct diagnoses from 75% to 92% in a randomised trial
EyeFM was tested as an assistant to 16 ophthalmologists screening 668 high-risk patients in China. Patients in the AI arm also followed referral advice more often.
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
FDA clears an AI tool that checks whether a feeding tube ended up in the wrong place
GE's Critical Care Suite gains an algorithm that flags misplaced enteric tubes on a chest X-ray — a complication that reviews of the practice have linked to respiratory harm and, in some cases, death.
China has approved 154 AI medical devices since 2020, and 69% of them read scans
A new audit of China's regulatory record finds a market concentrated in radiology, dominated by deep learning, and clustered in four cities. It also finds the approval curve flattening.