三模态蛋白质语言模型使高级蛋白质搜索成为可能。

IF 41.7 1区 生物学 Q1 BIOTECHNOLOGY & APPLIED MICROBIOLOGY
Jin Su,Yan He,Shiyang You,Shiyu Jiang,Xibin Zhou,Xuting Zhang,Yuxuan Wang,Xining Su,Igor Tolstoy,Xing Chang,Hongyuan Lu,Fajie Yuan
{"title":"三模态蛋白质语言模型使高级蛋白质搜索成为可能。","authors":"Jin Su,Yan He,Shiyang You,Shiyu Jiang,Xibin Zhou,Xuting Zhang,Yuxuan Wang,Xining Su,Igor Tolstoy,Xing Chang,Hongyuan Lu,Fajie Yuan","doi":"10.1038/s41587-025-02836-0","DOIUrl":null,"url":null,"abstract":"ProTrek unifies protein sequence, structure and natural language function in a trimodal language model through contrastive learning, enabling comprehensive searches between any two modalities, including within modality. ProTrek surpasses current alignment tools (for example, Foldseek and MMseqs2) in speed and accuracy for identifying functionally related proteins. Computational and wet-lab experimental validations show that the ProTrek server ( www.search-protrek.com ), with precomputed embeddings for over 5 billion proteins, efficiently processes and analyzes large-scale protein repositories.","PeriodicalId":19084,"journal":{"name":"Nature biotechnology","volume":"23 1","pages":""},"PeriodicalIF":41.7000,"publicationDate":"2025-10-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"A trimodal protein language model enables advanced protein searches.\",\"authors\":\"Jin Su,Yan He,Shiyang You,Shiyu Jiang,Xibin Zhou,Xuting Zhang,Yuxuan Wang,Xining Su,Igor Tolstoy,Xing Chang,Hongyuan Lu,Fajie Yuan\",\"doi\":\"10.1038/s41587-025-02836-0\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"ProTrek unifies protein sequence, structure and natural language function in a trimodal language model through contrastive learning, enabling comprehensive searches between any two modalities, including within modality. ProTrek surpasses current alignment tools (for example, Foldseek and MMseqs2) in speed and accuracy for identifying functionally related proteins. Computational and wet-lab experimental validations show that the ProTrek server ( www.search-protrek.com ), with precomputed embeddings for over 5 billion proteins, efficiently processes and analyzes large-scale protein repositories.\",\"PeriodicalId\":19084,\"journal\":{\"name\":\"Nature biotechnology\",\"volume\":\"23 1\",\"pages\":\"\"},\"PeriodicalIF\":41.7000,\"publicationDate\":\"2025-10-02\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Nature biotechnology\",\"FirstCategoryId\":\"5\",\"ListUrlMain\":\"https://doi.org/10.1038/s41587-025-02836-0\",\"RegionNum\":1,\"RegionCategory\":\"生物学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"BIOTECHNOLOGY & APPLIED MICROBIOLOGY\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Nature biotechnology","FirstCategoryId":"5","ListUrlMain":"https://doi.org/10.1038/s41587-025-02836-0","RegionNum":1,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"BIOTECHNOLOGY & APPLIED MICROBIOLOGY","Score":null,"Total":0}
引用次数: 0

摘要

ProTrek通过对比学习将蛋白质序列、结构和自然语言功能统一到一个三模态语言模型中,实现任意两模态之间的综合搜索,包括模态内部的搜索。在识别功能相关蛋白的速度和准确性方面,ProTrek超越了当前的比对工具(例如,Foldseek和MMseqs2)。计算和湿实验室实验验证表明,ProTrek服务器(www.search-protrek.com)预先计算了超过50亿个蛋白质的嵌入,可以有效地处理和分析大规模的蛋白质库。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
A trimodal protein language model enables advanced protein searches.
ProTrek unifies protein sequence, structure and natural language function in a trimodal language model through contrastive learning, enabling comprehensive searches between any two modalities, including within modality. ProTrek surpasses current alignment tools (for example, Foldseek and MMseqs2) in speed and accuracy for identifying functionally related proteins. Computational and wet-lab experimental validations show that the ProTrek server ( www.search-protrek.com ), with precomputed embeddings for over 5 billion proteins, efficiently processes and analyzes large-scale protein repositories.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
Nature biotechnology
Nature biotechnology 工程技术-生物工程与应用微生物
CiteScore
63.00
自引率
1.70%
发文量
382
审稿时长
3 months
期刊介绍: Nature Biotechnology is a monthly journal that focuses on the science and business of biotechnology. It covers a wide range of topics including technology/methodology advancements in the biological, biomedical, agricultural, and environmental sciences. The journal also explores the commercial, political, ethical, legal, and societal aspects of this research. The journal serves researchers by providing peer-reviewed research papers in the field of biotechnology. It also serves the business community by delivering news about research developments. This approach ensures that both the scientific and business communities are well-informed and able to stay up-to-date on the latest advancements and opportunities in the field. Some key areas of interest in which the journal actively seeks research papers include molecular engineering of nucleic acids and proteins, molecular therapy, large-scale biology, computational biology, regenerative medicine, imaging technology, analytical biotechnology, applied immunology, food and agricultural biotechnology, and environmental biotechnology. In summary, Nature Biotechnology is a comprehensive journal that covers both the scientific and business aspects of biotechnology. It strives to provide researchers with valuable research papers and news while also delivering important scientific advancements to the business community.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信