{"title":"Data Science at the Interface of Air Pollution and Lung Health: Toward Precision Health.","authors":"Kenneth Ofori-Amanfo, Maya Dustin, Emilia L Lim","doi":"10.1146/annurev-biodatasci-092724-061536","DOIUrl":"10.1146/annurev-biodatasci-092724-061536","url":null,"abstract":"<p><p>Air pollution is a leading cause of death, commonly linked to respiratory diseases, such as asthma, chronic obstructive pulmonary disease, and lung cancer. With a view toward precision health, efforts have been made to ascertain the relationships between air pollution and molecular dysregulation to direct risk management and prevention of pollution-promoted respiratory diseases. To this end, there have been several analyses aimed at uncovering how air pollution drives disease through dysregulation of the methylome, transcriptome, metabolome, proteome, genome, and microbiome. Here we review these studies, assess their current limitations, and discuss their contributions to research at the interface of air pollution and lung health. We highlight how large-scale analyses have elucidated the role of air pollution in dysregulating genomic stability, inflammation, and apoptotic pathways to promote respiratory diseases. Finally, we summarize opportunities for future research that may be facilitated by ongoing improvements in exposure estimates and multiomic integration strategies.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"141-170"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147783258","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Marina L Fernández, Felipe A Gajardo Escobar, Marcela K Sjöberg-Herrera, Pablo A Romagnoli, Danilo G Ceschin
{"title":"Beyond the Transcriptome: Leveraging Cellular Indexing of Transcriptomes and Epitopes by Sequencing (CITE-seq) for Deeper Cellular Insights.","authors":"Marina L Fernández, Felipe A Gajardo Escobar, Marcela K Sjöberg-Herrera, Pablo A Romagnoli, Danilo G Ceschin","doi":"10.1146/annurev-biodatasci-092724-061209","DOIUrl":"10.1146/annurev-biodatasci-092724-061209","url":null,"abstract":"<p><p>The advent of single-cell genomics has revealed profound cellular heterogeneity, yet transcriptomic profiling alone often fails to show the functional proteomic state of a cell. Cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) emerged to bridge this critical gap, enabling the simultaneous quantification of RNA and surface protein expression within individual cells. This transformative technology leverages oligonucleotide-conjugated antibodies to assess protein abundance using a panel of marker-specific tags, which are cocaptured with cellular messenger RNA in single-cell sequencing workflows. This review details the methodological principles of CITE-seq, its compatibility with diverse sequencing platforms, and the computational frameworks required for integrated data analysis. We highlight its impact on cell biology, oncology, and infectious disease research, where it has refined cell classification, dissected tumor microenvironments, and decoded host-pathogen interactions. Finally, we discuss persistent technical and computational challenges and outline future directions, including spatial integration and predictive modeling, positioning CITE-seq as a cornerstone of next-generation biomedical discovery.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"197-212"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147783339","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Artificial Intelligence in Image-Based Cardiovascular Disease Analysis.","authors":"Xin Wang, Mingcheng Hu, Connie W Tsao, Hongtu Zhu","doi":"10.1146/annurev-biodatasci-092624-111837","DOIUrl":"10.1146/annurev-biodatasci-092624-111837","url":null,"abstract":"<p><p>Recent advancements in artificial intelligence (AI) have significantly influenced the field of cardiovascular disease (CVD) analysis, particularly in image-based diagnostics. Our article presents an extensive review of AI applications in image-based CVD analysis, offering insights into its current state and future potential. We systematically categorize the literature based on the primary anatomical structures related to CVD, dividing them into nonvessel structures (such as ventricles and atria) and vessel structures (including the aorta and coronary arteries). This categorization provides a structured approach to explore various imaging modalities like computed tomography and magnetic resonance imaging, which are commonly used in CVD research. Our review encompasses these modalities, giving a broad perspective on the diverse imaging techniques integrated with AI for CVD analysis. We conclude with an examination of the challenges and limitations inherent in current AI-based CVD analysis methods and suggest directions for future research to overcome these hurdles.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"555-585"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148057212","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Ana Laura Hernández-Ledesma, Evelia Lorena Coss-Navarrete, Grecia Sevilla-Parra, María Fernanda Bravo-García, Alejandra Schäfer, Marion E G Brunck, Alejandra Medina-Rivera
{"title":"Exploring Genetic Variations Associated with the Immune Response in Underrepresented Populations.","authors":"Ana Laura Hernández-Ledesma, Evelia Lorena Coss-Navarrete, Grecia Sevilla-Parra, María Fernanda Bravo-García, Alejandra Schäfer, Marion E G Brunck, Alejandra Medina-Rivera","doi":"10.1146/annurev-biodatasci-092724-034109","DOIUrl":"10.1146/annurev-biodatasci-092724-034109","url":null,"abstract":"<p><p>The immune system protects the body through tightly coordinated innate and adaptive responses shaped largely by an individual's genetic background. Genetic variants associated with immune responses provide key pathways underlying disease susceptibility. Historically, the large majority of genomic studies have focused on European populations, potentially limiting our understanding of the relation between genetics and immune response in distinct ancestries. To assess this imbalance, we identified and characterized 206 studies from 370 recruitment sites worldwide that investigated immune-related traits in the GWAS Catalog. European cohorts dominated the data, with minimal representation from African, Latin American and Native American, Oceanian, and Asian ancestries. This imbalance limits discovery of ancestry-specific biology, reduces the accuracy of risk prediction outside Europe, and limits clinical translation. Expanding ancestry diversity in genomic research is essential to uncover the full spectrum of genetic architecture that shapes immune traits and diseases and to ensure that discoveries in immunogenomics benefit all populations on Earth.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"501-522"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148035734","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Manuel Corpas, Oyesola Ojewunmi, Heinner Guio, Segun Fatumo
{"title":"Electronic Health Record-Linked Biobank Expansion Reveals Global Health Inequities.","authors":"Manuel Corpas, Oyesola Ojewunmi, Heinner Guio, Segun Fatumo","doi":"10.1146/annurev-biodatasci-092724-030452","DOIUrl":"10.1146/annurev-biodatasci-092724-030452","url":null,"abstract":"<p><p>Electronic health record (EHR)-linked biobanks are transforming biomedical research, enabling population-scale studies that integrate genomic, clinical, and phenotypic data. Yet as these resources proliferate, it remains unclear how their research outputs reflect global health priorities. This article presents a comprehensive review of five globally established EHR-linked biobanks: UK Biobank, the Million Veteran Program, FinnGen, the All of Us Research Program, and the Estonian Biobank. Drawing on 14,142 peer-reviewed publications from 2000 to 2024, we show how each biobank displays a distinct thematic profile, shaped by institutional mandates, population focus, and methodological design. We further evaluate alignment with global disease burden by mapping biobank-linked publications to 25 high-priority disease areas using World Health Organization disability-adjusted life years data. Our burden-adjusted gap scores and opportunity indices reveal striking underrepresentation of conditions such as malaria, tuberculosis, and diarrheal diseases when comparing biobank research output against high-priority diseases.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"27-46"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147391363","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Privacy and Security Throughout the Health Data Life Cycle: From Primary Care to Research Networks.","authors":"Bradley A Malin, Chao Yan, Luca Bonomi","doi":"10.1146/annurev-biodatasci-092724-031932","DOIUrl":"10.1146/annurev-biodatasci-092724-031932","url":null,"abstract":"<p><p>Health data are increasingly generated, shared, and analyzed across an ever-growing collection of settings. While these developments enable new forms of biomedical discovery and clinical decision support, they also introduce evolving privacy, security, and trust challenges that extend beyond traditional regulatory and technical frameworks. In this review, we characterize the various risks and protections throughout the health data life cycle, from data generation and primary use in healthcare to secondary use in research and artificial intelligence (AI) model development. We discuss how regulation, organizational practices, and technological choices shape data protection requirements, and we discuss and contextualize emerging threats, such as incidental disclosures through AI tools. We further review technical approaches for mitigating these risks, including access control and auditing, re-identification risk assessment and statistical mechanisms for risk mitigation (e.g., differential privacy), and synthetic data generation. We also consider how collaboration across disparate organizations may be achieved through federated learning mechanisms and cryptographic technologies, such as secure multiparty computation. Throughout, we highlight trade-offs between privacy protection and data utility, and we articulate practical challenges in deploying these methods at scale. We conclude by identifying open issues for the field, including the need for standardized metrics and greater transparency to support trust in data-driven healthcare and research.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"47-67"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147515337","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Bradley Iott, Matthew Pantell, Julia Adler-Milstein
{"title":"Social Drivers of Health in the Electronic Health Record.","authors":"Bradley Iott, Matthew Pantell, Julia Adler-Milstein","doi":"10.1146/annurev-biodatasci-092724-023921","DOIUrl":"10.1146/annurev-biodatasci-092724-023921","url":null,"abstract":"<p><p>There is currently much interest in screening patients for social drivers of health (SDOHs) to address health-related social needs (HRSNs), prompting clinicians to increasingly document HRSNs in the electronic health record (EHR). SDOH data in the EHR create many opportunities to address patients' needs, tailor medical care to account for social circumstances, exchange these data across organizations, and enhance population health improvement efforts. This review synthesizes current efforts to document SDOH data in structured and unstructured methods, to make use of community-level SDOH data, and to develop national standards driving EHR SDOH documentation and exchange. We describe barriers faced by organizations related to data collection burden, patient privacy, data quality, and health equity, and we discuss future research priorities to make the collection and use of EHR SDOH data more accurate, actionable, and equitable.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"449-474"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148057188","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Modeling the Language of Codons with Artificial Intelligence.","authors":"Claudèle Lemay-St-Denis, Rachel Kolodny","doi":"10.1146/annurev-biodatasci-092624-105741","DOIUrl":"10.1146/annurev-biodatasci-092624-105741","url":null,"abstract":"<p><p>The messenger RNA (mRNA) sequence plays a central role in expressing functional proteins from genes. Species-specific nonuniform codon choices influence key processes, including mRNA stability, translation accuracy and efficiency, and cotranslational folding. As evolutionarily selected codon sequences encode multiple overlapping, yet often subtle, signals, advanced deep learning models are needed to identify their statistical patterns. This review summarizes recent progress in artificial intelligence (AI) models that operate at the codon level, including both discriminative and generative approaches. We discuss codon language models and their application to downstream prediction tasks, as well as generative models used to design codon sequences for heterologous expression and to recode proteins for improved mRNA stability and expression. Finally, we highlight studies leveraging codon-based AI models to gain insights into molecular evolution and outline several open problems in this rapidly developing field.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"475-500"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148035801","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Oliver J Bear Don't Walk Iv, Lauren W Yowelunh McLester-Davis, Dornell Pete, Danner Peter, Karina Hernandez-Hernandez, Alec J Calac, Valentín Q de la Sierra, Kelle Dhein, Kimiora Henare, Krystal S Tsosie
{"title":"Reclaiming Data, Restoring Health: The Indigenous Biomedical Data Science Renaissance.","authors":"Oliver J Bear Don't Walk Iv, Lauren W Yowelunh McLester-Davis, Dornell Pete, Danner Peter, Karina Hernandez-Hernandez, Alec J Calac, Valentín Q de la Sierra, Kelle Dhein, Kimiora Henare, Krystal S Tsosie","doi":"10.1146/annurev-biodatasci-092524-121351","DOIUrl":"10.1146/annurev-biodatasci-092524-121351","url":null,"abstract":"<p><p>Indigenous Peoples are leading a renaissance in biomedical data science. Throughout this narrative review, we highlight advancements from leaders in this renaissance across four phases of biomedical data science: ethics and values that inform research, data harmonization, model development and assessment, and commercialization. In addition, we synthesize five teachings that advance biomedical data science: Indigenous sovereignty, ethical stewardship, relationality and trust, community prioritization, and Indigenous worldviews. We conclude with considerations for all biomedical data scientists, Indigenous People in biomedical data science, and Indigenous leaders leveraging biomedical data science to advance Indigenous health and well-being. This renaissance offers a road map for global biomedical systems to evolve beyond colonial legacies toward practices that uphold self-determination, restore balance, embody relationality, and promote the well-being of human and nonhuman kin.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"523-553"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13349450/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148057129","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Genetic Modulators of Disease Penetrance.","authors":"Yunjun Kang, Theodore G Drivas","doi":"10.1146/annurev-biodatasci-103123-094547","DOIUrl":"10.1146/annurev-biodatasci-103123-094547","url":null,"abstract":"<p><p>In human genetics, penetrance describes the probability that a genotype manifests as a phenotype within a defined clinical window, yet its biological determinants remain poorly understood. Advances in sequencing technologies, statistical frameworks, and biobank-scale resources now enable systematic discovery of penetrance-modifying coding, regulatory, and structural variants spanning the allele frequency-effect size spectrum. Using metabolic dysfunction-associated steatotic liver disease, chronic kidney disease, and Alzheimer's disease as exemplars, we illustrate how rare, high-impact mutations, ancestry-enriched intermediate-frequency alleles, and common polygenic variation act through distinct molecular and cellular mechanisms to shape disease liability. We highlight gaps in integrating different variant types, modeling context-dependent effects, and building frameworks to translate genetic findings into individualized risk assessment. A unified understanding of penetrance promises to improve genotype-to-phenotype inference, refine patient risk stratification, and accelerate genetically informed therapeutic strategies.</p>","PeriodicalId":29775,"journal":{"name":"Annual Review of Biomedical Data Science","volume":" ","pages":"93-116"},"PeriodicalIF":8.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147700036","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}