Mitigating negative data bias to enhance TCR-epitope binding and residue interaction prediction.

IF 7.3 2区 生物学 Q1 BIOCHEMICAL RESEARCH METHODS
Xue Mi, Jinghua Zhu, Zhu Dai, Yuheng Zhu, Bo Ding, Hao Lin, Yang Shen, Guochun Cao, Zhongdang Xiao
{"title":"Mitigating negative data bias to enhance TCR-epitope binding and residue interaction prediction.","authors":"Xue Mi, Jinghua Zhu, Zhu Dai, Yuheng Zhu, Bo Ding, Hao Lin, Yang Shen, Guochun Cao, Zhongdang Xiao","doi":"10.1093/bib/bbag418","DOIUrl":null,"url":null,"abstract":"<p><p>Accurate prediction of the binding specificity between T-cell receptors (TCRs) and epitopes, along with the elucidation of their molecular interaction mechanisms, is pivotal for advancing immunotherapy and vaccine development. In this study, we propose a negative dataset construction strategy based on region-directed random mutations as an effective complement to traditional negative sampling methods. This strategy preserves the conserved amino acid motifs encoded by the V and J gene segments of the CDR3$\\beta$ sequence while introducing key residue mutations within the central junctional region. By constructing hard negatives, this approach encourages the model to capture more discriminative TCR-epitope binding features. Based on this optimized dataset, we developed TranTCR, a computational framework comprising two models: TranTCR-bind, which focuses on global sequence-level binding probability prediction, and TranTCR-map, which leverages transfer learning to translate global binding knowledge into fine-grained characterizations of residue-level interactions, such as inter-residue distances and contact scores. Experimental results demonstrate that TranTCR-bind exhibits superior predictive performance and generalization robustness across various negative sampling protocols. Furthermore, TranTCR-map utilizes attention mechanisms to deeply resolve complex inter-amino acid associations, enabling the identification of latent binding patterns and the revelation of TCR cross-reactivity characteristics. This study provides an efficient computational tool for the high-throughput screening of TCR repertoires and the digital characterization of immune recognition mechanisms.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3000,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13440130/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Briefings in bioinformatics","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1093/bib/bbag418","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"BIOCHEMICAL RESEARCH METHODS","Score":null,"Total":0}
引用次数: 0

Abstract

Accurate prediction of the binding specificity between T-cell receptors (TCRs) and epitopes, along with the elucidation of their molecular interaction mechanisms, is pivotal for advancing immunotherapy and vaccine development. In this study, we propose a negative dataset construction strategy based on region-directed random mutations as an effective complement to traditional negative sampling methods. This strategy preserves the conserved amino acid motifs encoded by the V and J gene segments of the CDR3$\beta$ sequence while introducing key residue mutations within the central junctional region. By constructing hard negatives, this approach encourages the model to capture more discriminative TCR-epitope binding features. Based on this optimized dataset, we developed TranTCR, a computational framework comprising two models: TranTCR-bind, which focuses on global sequence-level binding probability prediction, and TranTCR-map, which leverages transfer learning to translate global binding knowledge into fine-grained characterizations of residue-level interactions, such as inter-residue distances and contact scores. Experimental results demonstrate that TranTCR-bind exhibits superior predictive performance and generalization robustness across various negative sampling protocols. Furthermore, TranTCR-map utilizes attention mechanisms to deeply resolve complex inter-amino acid associations, enabling the identification of latent binding patterns and the revelation of TCR cross-reactivity characteristics. This study provides an efficient computational tool for the high-throughput screening of TCR repertoires and the digital characterization of immune recognition mechanisms.

减轻负数据偏差,增强tcr -表位结合和残基相互作用预测。
准确预测t细胞受体(TCRs)与表位之间的结合特异性,以及阐明它们的分子相互作用机制,对于推进免疫治疗和疫苗开发至关重要。在本研究中,我们提出了一种基于区域定向随机突变的负数据集构建策略,作为传统负采样方法的有效补充。该策略保留了CDR3$\beta$序列的V和J基因片段编码的保守氨基酸基序,同时在中心连接区引入关键残基突变。通过构建硬阴性,该方法鼓励模型捕获更多鉴别tcr -表位结合特征。基于此优化的数据集,我们开发了TranTCR,这是一个由两个模型组成的计算框架:TranTCR-bind和TranTCR-map,前者专注于全局序列级结合概率预测,后者利用迁移学习将全局结合知识转化为残基级相互作用的细粒度特征,如残基间距离和接触分数。实验结果表明,TranTCR-bind在各种负采样协议中表现出优异的预测性能和泛化鲁棒性。此外,TranTCR-map利用注意机制深入解析复杂的氨基酸间关联,从而识别潜在的结合模式,揭示TCR交叉反应特性。本研究为TCR谱的高通量筛选和免疫识别机制的数字化表征提供了一种高效的计算工具。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Briefings in bioinformatics
Briefings in bioinformatics 生物-生化研究方法
CiteScore
13.20
自引率
13.70%
发文量
549
审稿时长
6 months
期刊介绍: Briefings in Bioinformatics is an international journal serving as a platform for researchers and educators in the life sciences. It also appeals to mathematicians, statisticians, and computer scientists applying their expertise to biological challenges. The journal focuses on reviews tailored for users of databases and analytical tools in contemporary genetics, molecular and systems biology. It stands out by offering practical assistance and guidance to non-specialists in computerized methodologies. Covering a wide range from introductory concepts to specific protocols and analyses, the papers address bacterial, plant, fungal, animal, and human data. The journal's detailed subject areas include genetic studies of phenotypes and genotypes, mapping, DNA sequencing, expression profiling, gene expression studies, microarrays, alignment methods, protein profiles and HMMs, lipids, metabolic and signaling pathways, structure determination and function prediction, phylogenetic studies, and education and training.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书