Integrating machine learning and deep learning with multiple molecular fingerprints for topoisomerase I inhibitor screening and lead identification.

IF 4.3 2区 化学 Q2 CHEMISTRY, APPLIED
Huang Zeng, Xuerou Zheng, Jiayao Liu, Shengyuan Zhang, Hua Nie, Nan Wang, Lingfeng Wu, Chunfang Liu, Ming Zhai, Hao Yang, Jiunlong Yang, Bo Qiu
{"title":"Integrating machine learning and deep learning with multiple molecular fingerprints for topoisomerase I inhibitor screening and lead identification.","authors":"Huang Zeng, Xuerou Zheng, Jiayao Liu, Shengyuan Zhang, Hua Nie, Nan Wang, Lingfeng Wu, Chunfang Liu, Ming Zhai, Hao Yang, Jiunlong Yang, Bo Qiu","doi":"10.1007/s11030-026-11709-w","DOIUrl":null,"url":null,"abstract":"<p><p>Topoisomerase I (TOP1) is a crucial anticancer target, but the development of traditional TOP1 inhibitors suffers from long research cycles, high costs, and low success rates. Existing artificial intelligence (AI)-driven studies lack systematic comparisons of molecular fingerprints and algorithms, as well as user-friendly predictive application tools. To address these gaps, this study retrieved TOP1 inhibitor activity data from the ChEMBL database, integrated five types of molecular fingerprints (AtomPairs, MACCS, Morgan, PharmacoPFP, and RDKitDes), and constructed and compared classical machine learning (ML) models and deep learning (DL) models, resulting in a total of 40 models. The four top-performing models, SVM::Morgan, RF::Morgan, DNN::MACCS, and KNN::Morgan, achieved ROC-AUC values of 0.93-0.94 under random splitting. Y-scrambling supported that the models learned non-random structure-activity relationships, while SHAP analysis identified key molecular features. The URL of the developed web application is http://drugpred.top:5000 , and this application enables the prediction of TOP1 inhibitory activity via SMILES (Simplified Molecular-Input Line-Entry System) or molecular structure drawing. Additionally, standalone desktop applications (.exe) for offline prediction are freely available at https://github.com/zenghuang8006/TOP1-inhibitor-prediction . Screening of 189,554 SPECS compounds followed by in vitro validation identified AG60 and AI61 as potential TOP1 inhibitors hits, with inhibition rates of 64% and 90% at 400 µM, respectively. Overall, this study provides a practical computational framework for TOP1 inhibitor screening and identifies promising candidate compounds. Notably, scaffold-split AUC values decreased to 0.67-0.82, indicating reduced extrapolative performance for compounds containing previously unseen scaffolds.</p>","PeriodicalId":708,"journal":{"name":"Molecular Diversity","volume":" ","pages":""},"PeriodicalIF":4.3000,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Molecular Diversity","FirstCategoryId":"92","ListUrlMain":"https://doi.org/10.1007/s11030-026-11709-w","RegionNum":2,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"CHEMISTRY, APPLIED","Score":null,"Total":0}
引用次数: 0

Abstract

Topoisomerase I (TOP1) is a crucial anticancer target, but the development of traditional TOP1 inhibitors suffers from long research cycles, high costs, and low success rates. Existing artificial intelligence (AI)-driven studies lack systematic comparisons of molecular fingerprints and algorithms, as well as user-friendly predictive application tools. To address these gaps, this study retrieved TOP1 inhibitor activity data from the ChEMBL database, integrated five types of molecular fingerprints (AtomPairs, MACCS, Morgan, PharmacoPFP, and RDKitDes), and constructed and compared classical machine learning (ML) models and deep learning (DL) models, resulting in a total of 40 models. The four top-performing models, SVM::Morgan, RF::Morgan, DNN::MACCS, and KNN::Morgan, achieved ROC-AUC values of 0.93-0.94 under random splitting. Y-scrambling supported that the models learned non-random structure-activity relationships, while SHAP analysis identified key molecular features. The URL of the developed web application is http://drugpred.top:5000 , and this application enables the prediction of TOP1 inhibitory activity via SMILES (Simplified Molecular-Input Line-Entry System) or molecular structure drawing. Additionally, standalone desktop applications (.exe) for offline prediction are freely available at https://github.com/zenghuang8006/TOP1-inhibitor-prediction . Screening of 189,554 SPECS compounds followed by in vitro validation identified AG60 and AI61 as potential TOP1 inhibitors hits, with inhibition rates of 64% and 90% at 400 µM, respectively. Overall, this study provides a practical computational framework for TOP1 inhibitor screening and identifies promising candidate compounds. Notably, scaffold-split AUC values decreased to 0.67-0.82, indicating reduced extrapolative performance for compounds containing previously unseen scaffolds.

将机器学习和深度学习与多分子指纹相结合用于拓扑异构酶I抑制剂筛选和导联鉴定。
拓扑异构酶I (TOP1)是一种重要的抗癌靶点,但传统的TOP1抑制剂的开发存在研究周期长、成本高、成功率低等问题。现有的人工智能(AI)驱动的研究缺乏分子指纹和算法的系统比较,以及用户友好的预测应用工具。为了解决这些问题,本研究从ChEMBL数据库中检索了TOP1抑制剂的活性数据,整合了五种类型的分子指纹图谱(AtomPairs, MACCS, Morgan, PharmacoPFP和RDKitDes),并构建和比较了经典机器学习(ML)模型和深度学习(DL)模型,共建立了40个模型。SVM::Morgan、RF::Morgan、DNN::MACCS和KNN::Morgan这四个表现最好的模型在随机分裂下的ROC-AUC值为0.93-0.94。y - scramble支持模型学习非随机的结构-活性关系,而SHAP分析确定了关键的分子特征。开发的web应用程序的URL为http://drugpred.top:5000,该应用程序可以通过SMILES (Simplified molecular - input Line-Entry System)或分子结构图来预测TOP1的抑制活性。此外,用于离线预测的独立桌面应用程序(.exe)可以在https://github.com/zenghuang8006/TOP1-inhibitor-prediction上免费获得。通过筛选189,554个SPECS化合物并进行体外验证,发现AG60和AI61是潜在的TOP1抑制剂,在400µM时的抑制率分别为64%和90%。总的来说,本研究为TOP1抑制剂的筛选和确定有希望的候选化合物提供了一个实用的计算框架。值得注意的是,支架分裂的AUC值降至0.67-0.82,表明含有以前未见过的支架的化合物的外推性能降低。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Molecular Diversity
Molecular Diversity 化学-化学综合
CiteScore
7.30
自引率
7.90%
发文量
219
审稿时长
2.7 months
期刊介绍: Molecular Diversity is a new publication forum for the rapid publication of refereed papers dedicated to describing the development, application and theory of molecular diversity and combinatorial chemistry in basic and applied research and drug discovery. The journal publishes both short and full papers, perspectives, news and reviews dealing with all aspects of the generation of molecular diversity, application of diversity for screening against alternative targets of all types (biological, biophysical, technological), analysis of results obtained and their application in various scientific disciplines/approaches including: combinatorial chemistry and parallel synthesis; small molecule libraries; microwave synthesis; flow synthesis; fluorous synthesis; diversity oriented synthesis (DOS); nanoreactors; click chemistry; multiplex technologies; fragment- and ligand-based design; structure/function/SAR; computational chemistry and molecular design; chemoinformatics; screening techniques and screening interfaces; analytical and purification methods; robotics, automation and miniaturization; targeted libraries; display libraries; peptides and peptoids; proteins; oligonucleotides; carbohydrates; natural diversity; new methods of library formulation and deconvolution; directed evolution, origin of life and recombination; search techniques, landscapes, random chemistry and more;
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书