Maryam Yassi, Mark Ezegbogu, Euan J Rodger, Peter Stockwell, Aniruddha Chatterjee, Matthew Parry
{"title":"DNAmBERT: a transformer-based model for non-invasive cancer diagnosis using DNA sequence and methylation data.","authors":"Maryam Yassi, Mark Ezegbogu, Euan J Rodger, Peter Stockwell, Aniruddha Chatterjee, Matthew Parry","doi":"10.1093/bib/bbag455","DOIUrl":"10.1093/bib/bbag455","url":null,"abstract":"<p><p>DNA methylation alterations are early and stable hallmarks of cancer and represent promising biomarkers for non-invasive detection using circulating cell-free DNA (cfDNA). However, current computational approaches often model DNA sequence and methylation features separately and struggle to capture complex read-level methylation architecture in heterogeneous, low-signal liquid biopsy data. Here, we present DNAmBERT, a Transformer-based deep learning framework designed to jointly model DNA sequence context and read-level methylation haplotype structure from cfDNA methylation sequencing data. DNAmBERT integrates k-mer-encoded DNA sequences with methylation haplotype tokens using a unified representation and masked language modelling objective, enabling context-aware learning of sequence-epigenetic dependencies through self-attention. We evaluated DNAmBERT across multiple cfDNA methylation platforms (RRBS, cfRRBS, and cfMethyl-seq) and cancer types, including colorectal cancer, lung adenocarcinoma and hepatocellular carcinoma. In binary classification tasks, the model achieved high performance across platforms (AUC up to 0.99-1.00) and outperformed conventional machine learning and existing deep learning approaches. Aggregation of read-level predictions enabled quantitative tumour probability estimation at the sample level. Beyond binary detection, DNAmBERT supported multi-cancer and stage-aware classification, including early-stage disease, with multiclass AUC values up to 0.99. The framework further demonstrated effective cross-cancer transfer learning, maintaining robust performance under limited data availability. These results indicate that integrated sequence-haplotype representation learning provides an accurate and scalable approach for cfDNA-based multi-cancer detection.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13537528/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148878782","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Antonin Colajanni, Raluca Uricaru, Samuel Darko, Rahul Subramanian, Daniel C Douek, Rodolphe Thiébaut, Patricia Thebault
{"title":"Benchmarking methods for extracting microbial signal from host-dominated metatranscriptomes.","authors":"Antonin Colajanni, Raluca Uricaru, Samuel Darko, Rahul Subramanian, Daniel C Douek, Rodolphe Thiébaut, Patricia Thebault","doi":"10.1093/bib/bbag454","DOIUrl":"10.1093/bib/bbag454","url":null,"abstract":"<p><p>Human RNA sequencing (RNA-seq) data originally generated for human transcriptome profiling are overwhelmingly dominated by host sequences, yet they often contain a small fraction of non-human reads that can be exploited for microbial detection. When such datasets are repurposed for secondary microbiome-oriented analyses, extracting and accurately classifying this weak microbial signal becomes technically challenging, and no ready-to-use pipeline currently exists. In this study, we evaluate computational strategies for filtering host reads and classifying microbial transcripts in host-dominated RNA sequencing data. We compare assembly-based approaches similar to those used in a previous study focusing on microbial translocation with state-of-the-art assembly-free methods, and assess their respective strengths and limitations using simulated datasets reflecting low microbial abundance. Our results show that assembly-based methods yield accurate taxonomic predictions but struggle at low read depth, whereas assembly-free methods are more robust in sparse settings at the cost of reduced precision. To leverage the complementarity of both approaches, we propose a hybrid pipeline that integrates assembly-based and assembly-free classification. On simulated data, this hybrid strategy improves microbial classification performance compared with either approach alone. Application to a real human metatranscriptomic dataset analyzed in a microbial translocation context illustrates the broader microbial signal captured by the hybrid approach, despite intrinsic challenges related to the absence of reliable ground truth and the risk of host read misclassification. Our work provides a framework for extracting microbial signals from host-dominated human metatranscriptomes, enabling the reuse of existing transcriptomic datasets for microbiome-related analyses, including but not limited to microbial translocation studies.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13537457/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148878798","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Florencia R Díaz, Daniela Orschanski, Juan I Folco, Juana Espain Ceci, Guadalupe Nibeyro, Horderlin Robles Vega, Juan P Nicola, Elmer A Fernández
{"title":"Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics.","authors":"Florencia R Díaz, Daniela Orschanski, Juan I Folco, Juana Espain Ceci, Guadalupe Nibeyro, Horderlin Robles Vega, Juan P Nicola, Elmer A Fernández","doi":"10.1093/bib/bbag461","DOIUrl":"https://doi.org/10.1093/bib/bbag461","url":null,"abstract":"<p><p>Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning-based predictors such as AlphaFold2, OpenFold, and AlphaFold3 have transformed structural biology, enabling routine in silico modeling even for challenging or previously uncharacterized proteins. However, systematic benchmarking of these tools-especially for novel targets and single amino acid variants-remains limited. Conventional global metrics often fail to capture biologically meaningful discrepancies. By evaluating multiple implementations of AlphaFold2 and OpenFold, together with ColabFold and the AlphaFold3 server, across 10 different proteins and 222 single amino acid protein variants encompassing a wide range of sizes, structures, and functions, we show that although widely used global indicators-like mean pLDDT, pTM-score, and RMSD-frequently suggest comparable performance, substantial local-level differences remain elusive. To address this gap, we introduce a comparative framework leveraging Bland-Altman agreement analysis, to evaluate per-residue Cα-confidence differences and Per-Residue profiles (PRPs), complemented by Uniform Manifold Approximation and Projection (UMAP). This approach reveals marked localized divergences, particularly within flexible or intrinsically disordered regions, where both predictor choice and single-residue substitutions trigger the largest conformational shifts. We further demonstrate that using reduced homology databases has minimal impact on predicted structural quality, offering computationally efficient alternatives. Collectively, our findings underscore the importance of integrating global and residue-specific evaluations to more accurately assess robustness, agreement, and practical usability across contemporary protein structure prediction methods.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886372","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Correction to: scAnno: a deconvolution strategy-based automatic cell type annotation tool for single-cell RNA-sequencing data sets.","authors":"","doi":"10.1093/bib/bbag462","DOIUrl":"10.1093/bib/bbag462","url":null,"abstract":"","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13539269/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886320","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Olivér M Balogh, Mátyás Pétervári, Áron M Csernák, Eszter Puhl, András Horváth, Péter Ferdinandy, Bence Ágg
{"title":"Contrastive learning of adverse events to provide effective and interpretable vector representations for machine-assisted pharmacovigilance.","authors":"Olivér M Balogh, Mátyás Pétervári, Áron M Csernák, Eszter Puhl, András Horváth, Péter Ferdinandy, Bence Ágg","doi":"10.1093/bib/bbag463","DOIUrl":"https://doi.org/10.1093/bib/bbag463","url":null,"abstract":"<p><p>Post-marketing surveillance is crucial for drug safety, yet the tools of pharmacovigilance rely solely on text-based data that may limit contemporary machine learning methodologies in the support of decision-making. With the recent surge of employing large language models (LLMs) for text-based tasks, there also arises an unmet need for a different approach which is not grounded in the linguistic patterns of unfiltered natural text, like LLMs, but rather based on real-world drug safety data. Here, we adapt contrastive learning algorithms to generate adverse event vector representations from spontaneous adverse event reports to serve as machine-readable (i.e. numerical) resources for downstream pharmacovigilance applications, such as drug-event association prediction for signal detection or causality assessment. We present comprehensive interpretability analyses of the resulting representations through density-based clustering, semantic evaluation, and comparison of multivariate dispersions, revealing patterns that reflect both functional and causal relations of the adverse events while also capturing drug-safety-related information better than existing medical terminologies and encoder-only LLMs. Furthermore, we demonstrate the applicability of our representations as input features in our downstream classifier model, outperforming the reporting odds ratio method, commonly used by regulatory agencies, and also LLM-generated representations (area under the receiver operating characteristic curve: 0.88 versus 0.76-0.83) on drug-event association prediction benchmarks. Therefore, we propose an interpretable adverse event vector representation, serving as a general resource that could enable the development of a wide array of machine learning applications to support decision-making in pharmacovigilance and facilitate patient safety.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886328","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Jinbu Wang, Lili Du, Zhida Zhao, Li Qian, Keanning Li, Shiyuan Qiu, Meng Mao, Mang Liang, Zezhao Wang, Hongwei Li, Yan Chen, Bo Zhu, Caihong Zheng, Xue Gao, Lingyang Xu, Lupei Zhang, Junya Li, Huijiang Gao
{"title":"PGS-GS: a framework integrating polygenic scores and genomic selection in animal breeding.","authors":"Jinbu Wang, Lili Du, Zhida Zhao, Li Qian, Keanning Li, Shiyuan Qiu, Meng Mao, Mang Liang, Zezhao Wang, Hongwei Li, Yan Chen, Bo Zhu, Caihong Zheng, Xue Gao, Lingyang Xu, Lupei Zhang, Junya Li, Huijiang Gao","doi":"10.1093/bib/bbag397","DOIUrl":"10.1093/bib/bbag397","url":null,"abstract":"<p><p>Genomic prediction has become a central paradigm in biology, enabling quantitative inference of genetic contributions to complex traits across humans, animals, and plants. Although genomic research in human genetics and animal breeding shares a highly homologous methodological foundation, significant barriers persist in their analytical paradigms and application scenarios. This study aims to promote cross-disciplinary integration by introducing human-derived polygenic scores (PGS) algorithms into animal genomic selection (GS) and proposing a PGS-GS framework with a preliminary weighting-based implementation. We systematically benchmarked the predictive performance and computational efficiency of 20 algorithms, including classical linear models, machine learning, PGS, and PGS-GS using both array and whole-genome sequencing (WGS) data across four major agricultural species: beef cattle, sheep, pigs, and chickens. Our results demonstrate that PGS and PGS-GS algorithms achieve predictive accuracy competitive with genomic best linear unbiased prediction (GBLUP) while offering markedly higher computational efficiency. Moreover, incorporating PGS-derived prior information into weighted linear and non-linear models outperformed conventional weighted GBLUP. The results provide empirical evidence to inform algorithm selection and highlight the potential of integrating human-derived PGS methodologies into animal genomic prediction frameworks.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13533199/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148873146","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Gene-Chronos: parameter-efficient developmental time inference using a pretrained single-cell foundation model.","authors":"Yinbo Liu, Handi Gao, Tian Tian","doi":"10.1093/bib/bbag469","DOIUrl":"https://doi.org/10.1093/bib/bbag469","url":null,"abstract":"<p><p>Large-scale single-cell and spatial transcriptomic atlases enable the study of developmental processes at high resolution. However, most datasets capture only static snapshots of cells, making it difficult to infer continuous biological time from transcriptomic profiles. Existing temporal inference methods often show limited robustness across heterogeneous datasets, and recent single-cell foundation models, although powerful for representation learning, are not designed to capture continuous temporal relationships. We present Gene-Chronos, a parameter-efficient framework for developmental time inference built on a frozen pretrained Geneformer backbone. The model introduces learnable temporal prompt tokens and a temporal contrastive objective to extract time-informative signals and encourage temporally coherent organization of cell representations. Across multiple benchmark datasets spanning diverse species and developmental stages, Gene-Chronos outperforms existing approaches and demonstrates strong generalization to previously unseen samples. Attention-based analyses further identify genes associated with developmental progression, providing interpretable insights into temporal gene expression dynamics.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148890928","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Qi Wu, Yinbo Liu, Feng Yang, Weihong Huang, Xiaolei Zhu, Juan Liu
{"title":"Decoupling topological and molecular features for interpretable biomolecular interaction prediction.","authors":"Qi Wu, Yinbo Liu, Feng Yang, Weihong Huang, Xiaolei Zhu, Juan Liu","doi":"10.1093/bib/bbag471","DOIUrl":"https://doi.org/10.1093/bib/bbag471","url":null,"abstract":"<p><p>Predicting biomolecular interactions is fundamental to understanding cellular mechanisms and advancing drug discovery. However, biomolecular interactions exhibit immense diversity across multiple dimensions. Most existing computational methods are designed to handle one specific task or data modality, which limits their applicability and generalization capability in broader scenarios. To address this methodological rigidity, we propose a flexible framework for multi-modal feature fusion in biomolecular interaction prediction (FlexBIP). The core of FlexBIP lies in its modular architecture, which decouples intrinsic molecular features from complex graph topologies, enabling the adaptive integration of node attributes, edge properties, and auxiliary graph information. The flexible fusion methodology breaks through the limitations of task-specific models. This design enables FlexBIP to adaptively process and integrate biological data of different types and from various sources, including homogeneous interactions between molecules of the same type, heterogeneous interactions between different molecular classes, as well as qualitative binary, multi-class, and quantitative regression prediction tasks. Our research has yielded exciting results. In extensive testing across 15 benchmark datasets, covering 8 major categories of biomolecular associations, FlexBIP's performance comprehensively surpasses that of 25 state-of-the-art specialized models. Crucially, in data-scarce \"cold-start\" scenarios that simulate the discovery of new molecules, FlexBIP continues to demonstrate remarkable robustness and predictive accuracy. Furthermore, FlexBIP provides robust and reliable interpretability for various downstream analysis tasks.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148890939","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Correction to: Structure-informed machine learning for drug discovery: a task-centric perspective.","authors":"","doi":"10.1093/bib/bbag498","DOIUrl":"10.1093/bib/bbag498","url":null,"abstract":"","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13531227/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148863572","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Andriy Rebryk, Denice van Herwerden, Saer Samanipour, Peter Haglund
{"title":"Evaluation of peak alignment performance of commercial and open-source tools for GC×GC-MS based non-targeted screening.","authors":"Andriy Rebryk, Denice van Herwerden, Saer Samanipour, Peter Haglund","doi":"10.1093/bib/bbag458","DOIUrl":"10.1093/bib/bbag458","url":null,"abstract":"<p><p>Comprehensive two-dimensional gas chromatography-mass spectrometry (GC × GC-MS) is a powerful tool for analysing complex mixtures, but its wide use is limited by the lack of robust and transparent inter-sample peak alignment workflows. Commercial solutions often rely on proprietary algorithms, while user-friendly open-source alternatives remain scarce. Here, we introduce jAligner4GCxGC, a new open-access Julia package for inter-sample GC × GC-MS peak alignment and benchmark its performance within practical GC × GC-MS data processing workflows against Guineu (open-source) and ChromaTOF Sync 2D (commercial). The package performs data preprocessing, alignment, and postprocessing including feature merging, library searching, and compound data retrieval. Three indoor dust samples, including certified reference material, were spiked with >150 reference compounds at concentrations of 0.5, 5.0, and 50 pg/μl and analysed using workflows optimized for each application. Performance was evaluated based on recovery of correctly aligned spiked compounds and workflow processing time. Because the software packages differ in peak detection and preprocessing, comparisons involving Sync 2D should be interpreted at workflow rather than alignment algorithm level. At 50 pg/μl, Guineu and jAligner4GCxGC detected 94% and 92% of compounds, respectively, outperforming Sync 2D (82%). At 5 and 0.5 pg/μl, Sync 2D detected the highest proportions (79% and 71%), whereas Guineu and jAligner4GCxGC performed identically (68% and 33%). Alignment-related processing required seconds for Guineu, 4 min for jAligner4GCxGC, and 9 min for Sync 2D. However, Sync 2D provided the quickest end-to-end workflow. These results highlight practical trade-offs between commercial and open-source workflows and demonstrate that jAligner4GCxGC provides a transparent and reproducible solution for GC × GC-MS peak alignment.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13537445/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148878813","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}