{"title":"利用预测的表观基因组特征改进全基因组测序数据的多基因预测。","authors":"Wanwen Zeng, Hanmin Guo, Qiao Liu, Wing Hung Wong","doi":"10.1073/pnas.2419202122","DOIUrl":null,"url":null,"abstract":"<p><p>Polygenic risk scores (PRS) are essential tools for estimating individual susceptibility to complex diseases by aggregating the effects of many genetic variants. With the advent of whole-genome sequencing (WGS), rare and de novo variants can now be detected at scale, presenting new opportunities to enhance PRS performance. Additionally, regulatory mechanisms that govern gene expression play a critical role in disease manifestation, suggesting further potential for improvement. However, most existing PRS methods are not well-equipped to incorporate nonlinear variant effects, rare variant contributions, or regulatory context. To address these limitations, we developed Epi-PRS, a novel framework that leverages large language models (LLMs) to impute cell-type-specific epigenomic signals from personal diploid genotypes. These imputed signals act as informative intermediates between genotype and phenotype, allowing for more accurate modeling of variant impact. Our simulation studies demonstrate that Epi-PRS improves predictive accuracy by incorporating nonlinear relationships, rare variant effects, and regulatory information across large genomic regions. When applied to real data from the UK Biobank, Epi-PRS significantly outperforms existing PRS approaches in predicting risk for both breast cancer and type 2 diabetes. These results underscore the advantages of integrating WGS data, epigenomic context, and advanced LLMs framework to enhance both the predictive power and interpretability of PRS. Overall, Epi-PRS represents a promising step toward more precise and biologically informed disease risk prediction, with broad implications for advancing personalized medicine and understanding complex genetic architectures.</p>","PeriodicalId":20548,"journal":{"name":"Proceedings of the National Academy of Sciences of the United States of America","volume":"122 24","pages":"e2419202122"},"PeriodicalIF":9.1000,"publicationDate":"2025-06-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12184400/pdf/","citationCount":"0","resultStr":"{\"title\":\"Improving polygenic prediction from whole-genome sequencing data by leveraging predicted epigenomic features.\",\"authors\":\"Wanwen Zeng, Hanmin Guo, Qiao Liu, Wing Hung Wong\",\"doi\":\"10.1073/pnas.2419202122\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p><p>Polygenic risk scores (PRS) are essential tools for estimating individual susceptibility to complex diseases by aggregating the effects of many genetic variants. With the advent of whole-genome sequencing (WGS), rare and de novo variants can now be detected at scale, presenting new opportunities to enhance PRS performance. Additionally, regulatory mechanisms that govern gene expression play a critical role in disease manifestation, suggesting further potential for improvement. However, most existing PRS methods are not well-equipped to incorporate nonlinear variant effects, rare variant contributions, or regulatory context. To address these limitations, we developed Epi-PRS, a novel framework that leverages large language models (LLMs) to impute cell-type-specific epigenomic signals from personal diploid genotypes. These imputed signals act as informative intermediates between genotype and phenotype, allowing for more accurate modeling of variant impact. Our simulation studies demonstrate that Epi-PRS improves predictive accuracy by incorporating nonlinear relationships, rare variant effects, and regulatory information across large genomic regions. When applied to real data from the UK Biobank, Epi-PRS significantly outperforms existing PRS approaches in predicting risk for both breast cancer and type 2 diabetes. These results underscore the advantages of integrating WGS data, epigenomic context, and advanced LLMs framework to enhance both the predictive power and interpretability of PRS. Overall, Epi-PRS represents a promising step toward more precise and biologically informed disease risk prediction, with broad implications for advancing personalized medicine and understanding complex genetic architectures.</p>\",\"PeriodicalId\":20548,\"journal\":{\"name\":\"Proceedings of the National Academy of Sciences of the United States of America\",\"volume\":\"122 24\",\"pages\":\"e2419202122\"},\"PeriodicalIF\":9.1000,\"publicationDate\":\"2025-06-17\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12184400/pdf/\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the National Academy of Sciences of the United States of America\",\"FirstCategoryId\":\"103\",\"ListUrlMain\":\"https://doi.org/10.1073/pnas.2419202122\",\"RegionNum\":1,\"RegionCategory\":\"综合性期刊\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"2025/6/12 0:00:00\",\"PubModel\":\"Epub\",\"JCR\":\"Q1\",\"JCRName\":\"MULTIDISCIPLINARY SCIENCES\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the National Academy of Sciences of the United States of America","FirstCategoryId":"103","ListUrlMain":"https://doi.org/10.1073/pnas.2419202122","RegionNum":1,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/6/12 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"MULTIDISCIPLINARY SCIENCES","Score":null,"Total":0}
Improving polygenic prediction from whole-genome sequencing data by leveraging predicted epigenomic features.
Polygenic risk scores (PRS) are essential tools for estimating individual susceptibility to complex diseases by aggregating the effects of many genetic variants. With the advent of whole-genome sequencing (WGS), rare and de novo variants can now be detected at scale, presenting new opportunities to enhance PRS performance. Additionally, regulatory mechanisms that govern gene expression play a critical role in disease manifestation, suggesting further potential for improvement. However, most existing PRS methods are not well-equipped to incorporate nonlinear variant effects, rare variant contributions, or regulatory context. To address these limitations, we developed Epi-PRS, a novel framework that leverages large language models (LLMs) to impute cell-type-specific epigenomic signals from personal diploid genotypes. These imputed signals act as informative intermediates between genotype and phenotype, allowing for more accurate modeling of variant impact. Our simulation studies demonstrate that Epi-PRS improves predictive accuracy by incorporating nonlinear relationships, rare variant effects, and regulatory information across large genomic regions. When applied to real data from the UK Biobank, Epi-PRS significantly outperforms existing PRS approaches in predicting risk for both breast cancer and type 2 diabetes. These results underscore the advantages of integrating WGS data, epigenomic context, and advanced LLMs framework to enhance both the predictive power and interpretability of PRS. Overall, Epi-PRS represents a promising step toward more precise and biologically informed disease risk prediction, with broad implications for advancing personalized medicine and understanding complex genetic architectures.
期刊介绍:
The Proceedings of the National Academy of Sciences (PNAS), a peer-reviewed journal of the National Academy of Sciences (NAS), serves as an authoritative source for high-impact, original research across the biological, physical, and social sciences. With a global scope, the journal welcomes submissions from researchers worldwide, making it an inclusive platform for advancing scientific knowledge.