{"title":"Expression of Concern: Advances and challenges in single-cell RNA sequencing data analysis: a comprehensive review.","authors":"","doi":"10.1093/bib/bbag433","DOIUrl":"10.1093/bib/bbag433","url":null,"abstract":"","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13393891/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148577211","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Liye Zhang, Ran Yan, Weiming Gong, Xiang Zhou, Lu Liu, Zhongshang Yuan
{"title":"METEOR: a data-adaptive Mendelian randomization method for powerful detection of shared and specific exposures underlying multiple outcomes.","authors":"Liye Zhang, Ran Yan, Weiming Gong, Xiang Zhou, Lu Liu, Zhongshang Yuan","doi":"10.1093/bib/bbag364","DOIUrl":"10.1093/bib/bbag364","url":null,"abstract":"<p><p>Accurate identification of causal exposures for multimorbidity can benefit the co-prevention and co-management of multiple-related outcomes. This goal can be conceptually addressed within a multi-outcome Mendelian randomization (MR) framework. However, existing multi-outcome MR methods suffer from restrictions on format and availability of data inputs, fail to account for the potential sample overlap, rely on pre-selected independent instrumental variables (IVs), and are unable to account for horizontal pleiotropy. Here, we propose METEOR, a novel MR method that jointly models one exposure and multiple outcomes to identify both shared and outcome-specific causal exposures. METEOR accounts for sample overlap between exposure and outcomes, allows outcomes from different genome-wide association studies (GWAS) datasets, self-adaptively determines IVs from correlated single-nucleotide polymorphisms, and explicitly models horizontal pleiotropy. Using summary statistics, METEOR infers causal effects under a joint-likelihood framework with a scalable, sampling-based algorithm. Simulations show that METEOR presents well-calibrated $P$-values for both global and single-outcome tests, and achieves average power improvements of 55.33% and 56.50% over five existing MR methods in the global and single tests, respectively. In real data applications, METEOR produces the most accurate causal effect estimates in positive control analyses, reduces false positives by 18.75% in negative control analyses, and highlights that controlling BMI could benefit the co-management of multiple cardiovascular diseases (CVDs) and multiple gastrointestinal (GI) diseases, while controlling blood pressure could benefit the co-management of multimorbidity across CVDs and mental disorders (MDs), as well as across GI diseases and MDs.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13336660/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148396086","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"T-cell receptor alpha to beta chains binding prediction.","authors":"Devora Siminovsky, Yoram Louzoun","doi":"10.1093/bib/bbag358","DOIUrl":"10.1093/bib/bbag358","url":null,"abstract":"<p><p>Binding of T-cell receptors (TCRs) and their cognate peptide-major histocompatibility complex (pMHC) target is determined by both TCR$alpha $ and TCR$beta $ chains. However, not all TCR$alpha $ and TCR$beta $ can bind to each other. Predicting their pairing is crucial for understanding the TCR-pMHC interaction and developing effective de novo TCRs. Here, we show that in the general TCR repertoire, TCR$alpha $ and TCR$beta $ chain compositions are independent. However, in pMHC-binding TCRs, clear associations between TCR$alpha $ and TCR$beta $ chains are found, also for TCRs binding to the same pMHC. The association between the CDR3 amino acid composition and $V$, $J$ usage of TCR$alpha $ and TCR$beta $ reveals distinct binding patterns between specific $V$ and $J$ genes, as well as negative correlations between the charge and polarity of the TCR$alpha $ and TCR$beta $ chains, but positive associations between their molecular weights. These associations are used for the development of a prediction model for TCR$alpha $ and TCR$beta $ pairing. We present here TCR-BARN (TCR Beta-Alpha chains paiRing using Nlp) that employs an initial embedding for each amino acid in the TCR alpha and beta CDR3 sequences, followed by long short-term memory (LSTM) networks to capture sequence dependencies. The $V$ and $J$ genes are represented using one-hot encoding. LSTM outputs are concatenated and passed through a fully connected feedforward layer for binding prediction. TCR-BARN reaches an area under the curve $>0.65pm 0.007$ for epitope-bound TCRs. TCR-BARN can be used for generating cognate TCRs resembling natural TCRs and evaluating the generated TCR quality.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13345391/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148410185","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"AbTune: layer-wise selective fine-tuning of protein language models for antibodies.","authors":"Xiaotong Xu, Alexandre M J J Bonvin","doi":"10.1093/bib/bbag374","DOIUrl":"10.1093/bib/bbag374","url":null,"abstract":"<p><p>Antibodies play central roles in immune defense and are widely used as therapeutic agents. However, the high structural and sequence diversity of antigen-binding loops, combined with limited experimental data and weak co-evolutionary signals, makes it difficult to develop generalizable predictive models. In this work, we investigate test-time fine-tuning strategies to improve protein language model (pLM) performance in low-data settings, with a focus on antibody-related tasks. Systematic evaluations across tasks show that carefully constrained fine-tuning greatly enhances performance while preserving generalization. In particular, depth-selective fine-tuning consistently outperforms full-depth fine-tuning, with optimal performance achieved when tuning 50%-75% of model layers for medium- to small-sized pLMs. We introduce AbTune, a test-time fine-tuning framework that leverages this depth-controlled adaptation strategy. Across antibody structure prediction, mutation effect prediction, and binding affinity prediction, AbTune outperforms both standard pLM baselines and task-specific predictors, achieving the best performance among the evaluated baselines on two of the three tasks. To gain insight into the adaptation process and identify optimal AbTune protocols, we analyzed representation shifts, examined how sequence properties influence fine-tuning dynamics, and evaluated metrics that capture potential overfitting. Our results show that fine-tuning depth, duration, and perplexity jointly influence performance and must be carefully controlled to achieve optimal results.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13356902/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148429822","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Zülal Bingöl, Berkan Şahin, Klea Zambaku, Ricardo Roman-Brenes, Konstantina Koliogeorgi, Can Firtina, Onur Mutlu, Can Alkan
{"title":"De Bruijn graphs for pangenomics: in-depth performance benchmarking of de Bruijn graph-based tools for read mapping.","authors":"Zülal Bingöl, Berkan Şahin, Klea Zambaku, Ricardo Roman-Brenes, Konstantina Koliogeorgi, Can Firtina, Onur Mutlu, Can Alkan","doi":"10.1093/bib/bbag440","DOIUrl":"10.1093/bib/bbag440","url":null,"abstract":"<p><p>De Bruijn graphs are widely used in pangenome representation due to their numerous advantages and extensions, such as colored and compacted variants that enhance the representation of genetic variation. Although de Bruijn graphs are becoming increasingly adopted, their performance and energy impact have not been clearly studied. Such an overlooked understanding can lead to suboptimal designs for de Bruijn graph-based tools in addressing the computational challenges posed by pangenome data. To identify workflow bottlenecks and assess the efficiency of hardware utilization, we present an in-depth performance analysis of state-of-the-art de Bruijn graph-based read mapping tools on pangenomic datasets, focusing on scalability of execution time, hardware resource utilization, and energy consumption. We observe that the tools primarily prioritize data parallelism for processing read datasets, disregarding the increasing complexity of the pangenome graph, which hinders scalability. As the pangenome graph grows in size and complexity, cache miss rates also increase, leading to poor overall performance. By extensively analyzing sources of suboptimal performance, we pave the way for optimizing the existing and future tools to fully realize their potential in advancing pangenome research.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13499456/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148788509","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pavel Kohout, David Lacko, Milos Musil, Simeon Borko, Martin Stepanek, Jan Velecky, Petr Kabourek, Rayyan Tariq Khan, Monika Rosinska, Jiri Damborsky, Stanislav Mazurenko, David Bednar
{"title":"FireProtASR 2.0: evolution-guided Design of Protein Ancestors and Successors with phylogenetics and machine learning.","authors":"Pavel Kohout, David Lacko, Milos Musil, Simeon Borko, Martin Stepanek, Jan Velecky, Petr Kabourek, Rayyan Tariq Khan, Monika Rosinska, Jiri Damborsky, Stanislav Mazurenko, David Bednar","doi":"10.1093/bib/bbag436","DOIUrl":"10.1093/bib/bbag436","url":null,"abstract":"<p><p>Evolution-guided protein design remains one of the most effective strategies for engineering proteins with enhanced stability, activity, or specificity. To make these approaches more accessible, we previously developed FireProtASR-a fully automated pipeline for ancestral sequence reconstruction (ASR). Here, we present FireProtASR 2.0, a significantly enhanced version that extends the design space beyond ancestral inference by integrating a successor sequence predictor (SSP) and a generative model based on variational autoencoders (VAEs). These new modules enable both 'prospective' and 'retrospective' evolutionary design strategies. The SSP module predicts likely future mutations based on site-wise evolutionary trends, and the method was previously validated through in silico benchmarks, demonstrating improvements in thermostability and activity. The VAE module captures global evolutionary constraints in a low-dimensional latent space, from which novel functional ancestral-like variants can be sampled. The VAE-based design strategy was previously validated experimentally on the haloalkane dehalogenase family, yielding variants with enhanced thermostability while maintaining catalytic activity. Both these modules are newly available in FireProtASR in a fully automated pipeline, guiding the users via an interactive graphical user interface. With expanded functionality, modernized user interface, and a more robust backend, FireProtASR 2.0 provides a comprehensive, accessible, and fully automated platform for evolutionary-based protein engineering (https://loschmidt.chemi.muni.cz/fireprotasr/).</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518065/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148825410","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Xue Mi, Jinghua Zhu, Zhu Dai, Yuheng Zhu, Bo Ding, Hao Lin, Yang Shen, Guochun Cao, Zhongdang Xiao
{"title":"Mitigating negative data bias to enhance TCR-epitope binding and residue interaction prediction.","authors":"Xue Mi, Jinghua Zhu, Zhu Dai, Yuheng Zhu, Bo Ding, Hao Lin, Yang Shen, Guochun Cao, Zhongdang Xiao","doi":"10.1093/bib/bbag418","DOIUrl":"10.1093/bib/bbag418","url":null,"abstract":"<p><p>Accurate prediction of the binding specificity between T-cell receptors (TCRs) and epitopes, along with the elucidation of their molecular interaction mechanisms, is pivotal for advancing immunotherapy and vaccine development. In this study, we propose a negative dataset construction strategy based on region-directed random mutations as an effective complement to traditional negative sampling methods. This strategy preserves the conserved amino acid motifs encoded by the V and J gene segments of the CDR3$beta$ sequence while introducing key residue mutations within the central junctional region. By constructing hard negatives, this approach encourages the model to capture more discriminative TCR-epitope binding features. Based on this optimized dataset, we developed TranTCR, a computational framework comprising two models: TranTCR-bind, which focuses on global sequence-level binding probability prediction, and TranTCR-map, which leverages transfer learning to translate global binding knowledge into fine-grained characterizations of residue-level interactions, such as inter-residue distances and contact scores. Experimental results demonstrate that TranTCR-bind exhibits superior predictive performance and generalization robustness across various negative sampling protocols. Furthermore, TranTCR-map utilizes attention mechanisms to deeply resolve complex inter-amino acid associations, enabling the identification of latent binding patterns and the revelation of TCR cross-reactivity characteristics. This study provides an efficient computational tool for the high-throughput screening of TCR repertoires and the digital characterization of immune recognition mechanisms.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13440130/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148677289","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"RAG: a regularized adaptive graph-based method for rare-cell identification from single-cell expression data.","authors":"Xingsu Wang, Yanyan Chen, Dian Huang, Zhen Ju, Qi Wei, Shu Li, Shengzhong Feng","doi":"10.1093/bib/bbag379","DOIUrl":"10.1093/bib/bbag379","url":null,"abstract":"<p><p>Rare-cell identification is essential for dissecting disease mechanisms and developmental programs. Existing methods mostly rely on fixed-size neighbourhood graphs to separate rare-cell populations in single-cell expression data, which may embed rare cells into dominant clusters under varying sampling densities. This paper proposes the RAG method for identifying rare cells based on regularized adaptive graphs, which can better separate rare cells. Specifically, the regularized adaptive graph is constructed by estimating cell-specific radii from Euclidean-cosine hybrid dissimilarity to constrain effective neighbours and stabilize the adjacency, and then, assigning locally scaled hybrid affinities to make affinity magnitudes comparable across density-varying regions. Across 10 real single-cell RNA sequencing datasets, RAG overall outperformed six state-of-the-art methods, improving precision, F1 score, and rare-type coverage rate over the second-ranked baseline by 42%, 26%, and 35%, respectively. A case study on colorectal tumour tissue shows that RAG is more accurate in recovering annotated rare-cell populations and separating the substructure from the major population than the other evaluated methods. Further analyses on mouse airway epithelium and two pancreas datasets showed that about half of RAG-resolved small clusters corresponded to known annotated populations or marker-supported subpopulations. The source code is available at https://github.com/wangxingsu/RAG.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13379077/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148469187","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Richard Hillis, Nadya B A Johari, Masoud Shirali, Ian M Overton
{"title":"Data-intensive immune network modelling for One Health.","authors":"Richard Hillis, Nadya B A Johari, Masoud Shirali, Ian M Overton","doi":"10.1093/bib/bbag420","DOIUrl":"10.1093/bib/bbag420","url":null,"abstract":"<p><p>The immune system plays a critical role in morbidity and mortality. For example, infectious disease, cancer, and autoimmune disorders impose substantial health burdens upon both humans and animals. Immune systems share fundamental organisation, and the transmission of pathogens across species boundaries means that immune health in one population shapes risk in others. The One Health approach recognises this interdependence as the basis for understanding human, animal, and environmental health together. Immune systems are formed from multiscale networks of biomolecular interactions spanning specialised cell types, dynamic states, and differentiation trajectories. Data-intensive computational methods are required to model these processes accurately in specific biological contexts. Accordingly, bioinformatics is essential for understanding how the immune system functions under various conditions that arise from infections, in chronic diseases, and through environmental exposures. This article reviews cutting-edge techniques for studying immune function in human and animal health; with emphasis upon genetics, transcriptomics, single-cell approaches, network biology, and machine learning. We consider bioinformatics applications that inform our understanding of immune function to improve health and food systems. Examples are discussed from the rapidly developing cross-disciplinary landscape of computational and physical techniques. We illustrate data-intensive approaches in understanding context-specific immune biology, applied to illuminate the relationship between genetic variation and disease phenotypes.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13446512/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148683275","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Hanshi Xu, Guangquan Zhang, Hua Lin, Mark Grosser, Jie Lu
{"title":"WDCN: a comprehensive neural network based approach for estimating breast cancer risk.","authors":"Hanshi Xu, Guangquan Zhang, Hua Lin, Mark Grosser, Jie Lu","doi":"10.1093/bib/bbag412","DOIUrl":"10.1093/bib/bbag412","url":null,"abstract":"<p><p>Breast cancer is one of the most distressing cancers affecting women, and early detection is considered the most effective way to reduce breast cancer mortality. However, the benefits of early detection vary among different risk groups. Therefore, using a combination of genetic information, family history, and other factors to stratify populations by risk can help more people benefit from early detection. Traditional polygenic risk score (PRS) is essentially a weighted sum calculation method that has achieved some success, but it neglects the interactions between genes-genes, genes-environment, and their potential impact on breast cancer risk. In this context, we developed a new deep learning-based method called wide, deep, and cross network (WDCN). Experimental results show that our algorithm outperforms PRS and other machine learning baseline methods and achieves an area under the receiver operating characteristic curve (AUROC) of 0.6439 when using 286 single nucleotide polymorphism (SNP) features and 0.8865 when incorporating environmental features with genetic data. Increasing the SNP set to 317 further raised the performance to 0.6464 and 0.8872, both with and without non-genetic factors. Risk stratification shows that individuals in the top 30% have a relative risk of 7.85 (95% CI: 6.98-8.83) compared with those in the bottom 30%. We also identified an interaction between rs2588809 and age. This novel approach has shown promise for initial risk stratification of populations, potentially providing better decision-making support for individuals and clinicians.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13446515/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148683388","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}