Diverse genome data improved polygenic risk scores — but not evenly across traits
Using 245,388 genomes from the US All of Us programme, researchers built risk scores for 32 traits. More diversity helped prediction in under-represented groups — but a bigger sample was not always better.
Adding a large, ancestrally diverse dataset to the material used to build polygenic risk scores improved how well those scores predicted disease in under-represented groups — but the gain depended on the trait, and simply piling on more data was not always the best strategy [s1]. That is the headline finding of a study using 245,388 whole-genome sequences from the US All of Us research programme alongside UK Biobank data to construct multiancestry scores for 32 traits and diseases [s1].
What a polygenic score is, and why ancestry matters
A polygenic risk score adds up the small effects of many common genetic variants into a single number meant to estimate a person's inherited liability for a trait or disease. The catch, documented for years, is that these scores travel badly. Because the genome-wide association studies they are trained on have been dominated by people of European ancestry, the resulting scores predict less accurately in people whose ancestry differs from the training data — a portability problem that a widely cited 2019 analysis warned could, if scores were deployed as they stood, widen rather than narrow health disparities [s2].
The mechanism is not that other populations are "less genetic." It is statistical: the frequency of variants, and the way nearby variants are correlated, differ between populations, so a score calibrated in one group loses precision in another [s2]. The proposed fix has been to build the underlying association studies from more diverse participants. All of Us, which recruited heavily from groups long absent from genomic research, is a test of whether that fix works in practice.
What the study did
Researchers combined 245,388 whole-genome sequences from All of Us with UK Biobank data and developed multiancestry polygenic scores for 32 traits and diseases [s1]. They then examined how three things shaped a score's accuracy: the ancestry of the person being scored, the statistical method used to build the score, and the genetic architecture of the trait — that is, whether a trait is influenced by a very large number of variants of tiny effect or a smaller number of larger ones [s1].
Two design choices are worth stating plainly. First, this is a study of prediction accuracy — how closely a score tracks the trait — not of clinical outcomes. It does not show that using these scores in care changes what happens to patients. Second, the scores were evaluated across a spectrum of ancestry rather than in tidy racial categories, which is closer to how human genetic variation actually distributes.
The findings, and the twist
Increasing the diversity of the training data did improve accuracy for several traits, and the improvement was most pronounced in under-represented populations [s1]. That is the result the diverse-biobank strategy was built to produce, and it held.
The twist is that more data was not a universal good. Maximising sample size by meta-analysing All of Us and UK Biobank together — the obvious move — was not always optimal [s1]. For less polygenic traits, training on All of Us alone performed best in participants of African ancestry, a pattern the authors attribute to ancestry-enriched genetic effects that a larger, mostly European dataset can dilute rather than sharpen [s1].
The study also quantified the portability problem as a gradient rather than a cliff: an individual's score accuracy declined roughly linearly as their ancestry diverged from that of the discovery association study [s1]. Crucially, that decay was attenuated — flattened, not eliminated — when the score was trained on multiancestry data [s1].
What it does and does not establish
The honest reading is that representative biobanks measurably help, and help most where the need is greatest, but that "add more data" is too blunt a rule. Which data, for which trait, evaluated in whom, all matter — and for some traits the widely assumed strategy of pooling the biggest cohorts is the wrong call [s1].
Several limits bound the claim. The scores were assessed for predictive accuracy, and better accuracy is a precondition for clinical usefulness, not a demonstration of it; whether acting on a score improves health is a separate question that trials answer. Accuracy also remained sensitive to ancestral distance even after multiancestry training, so the equity gap was narrowed, not closed [s1]. And a score is a population-level statistic applied to an individual: it estimates relative liability, not destiny.
Why it matters
Polygenic scores are moving from research into pilots for conditions from breast cancer to coronary disease, and the risk that they perform worse for exactly the populations already underserved by medicine is not hypothetical [s2]. This study is evidence that the diverse-data remedy works — with the important qualification that it must be applied trait by trait rather than as a blanket rule. It strengthens the case for continued investment in representative genomic datasets, which is the concrete lever regulators and funders actually control.
The broader field is grappling with the same portability question in different populations: Taiwan's precision-medicine effort has tested scores built for Han Chinese participants, and population-specific genetic signals have turned up in work such as the genomics of high-altitude adaptation in the Andes. How such scores should be explained to the people who receive them is its own unsettled problem, where the trial evidence on risk communication has so far been thin.
What to watch
The next questions are whether these accuracy gains survive into clinical outcomes, whether the trait-by-trait pattern generalises to other diverse biobanks, and whether the tooling to pick the right training strategy per trait can be standardised. Until then, the practical message for anyone offered a polygenic score is that its reliability still depends partly on how closely their ancestry matches the data it was built from [s1].
Sources
- [s1] All of Us diversity and scale yield context-dependent improvements in polygenic prediction, Nature Genetics, 14 September 2026. https://doi.org/10.1038/s41588-026-02734-4
- [s2] Clinical use of current polygenic risk scores may exacerbate health disparities, Nature Genetics, 29 March 2019. https://doi.org/10.1038/s41588-019-0379-x
Sources
- All of Us diversity and scale yield context-dependent improvements in polygenic prediction — Nature Genetics , September 14, 2026
- Clinical use of current polygenic risk scores may exacerbate health disparities — Nature Genetics , March 29, 2019
Telling people their polygenic risk score changes almost nothing
A meta-analysis of 27 randomised trials found that disclosing a polygenic risk score did not meaningfully shift diet, screening uptake, medication use, anxiety or cholesterol.
A 565,390-person Taiwanese cohort tests polygenic scores outside European ancestry
Two Nature papers describe the Taiwan Precision Medicine Initiative and the risk scores built from it. The genetic effects identified explained up to 10.3% of health variation in the cohort.
Dubai's premarital genomic screening flagged 8% of couples as at risk
A mandatory citywide programme sequenced 782 genes in 1,000 prospective couples. The at-risk rate was double that of a comparable Australian study, and most flagged couples still chose to marry.
An Israeli HMO biobank sequenced 1,038 patients to hunt for deafness genes
Linking exome data to electronic medical records solved 15% of unexplained hearing-loss cases and flagged new candidate genes — while showing what the records could not supply.