Social and biological factors influence health outcomes through conventional pathways and, increasingly, through digital pathways of data and algorithms.
Biomedical data disadvantage is a health risk factor for most of the world's population; representation bias and distribution shifts in health data are key mechanisms through which it acts.
AI greatly empowers precision medicine but simultaneously opens up a major pathway through which this risk factor can exert its effects.
AI-powered precision medicine can be systematically less precise for data-disadvantaged populations, producing new or amplified health disparities. These disparities may vary continuously with genetic ancestry rather than conforming to discrete population categories.
These health disparities can affect any disease for which biomedical data inequality exists, making their potential impacts broad.
The digital pathways also provide targets for algorithmic intervention. Building on transfer learning, we are advancing toward in-context learning (ICL) with tabular foundation models and moving from discrete ancestry categories to the genetic ancestry continuum to develop a new framework for robust omics-based prediction of cancer risk and outcomes, aiming for a Pareto improvement: improving prediction for data-disadvantaged populations without reducing performance for populations already well served by existing models.