Huang Zeng, Xuerou Zheng, Jiayao Liu, Shengyuan Zhang, Hua Nie, Nan Wang, Lingfeng Wu, Chunfang Liu, Ming Zhai, Hao Yang, Jiunlong Yang, Bo Qiu
{"title":"Integrating machine learning and deep learning with multiple molecular fingerprints for topoisomerase I inhibitor screening and lead identification.","authors":"Huang Zeng, Xuerou Zheng, Jiayao Liu, Shengyuan Zhang, Hua Nie, Nan Wang, Lingfeng Wu, Chunfang Liu, Ming Zhai, Hao Yang, Jiunlong Yang, Bo Qiu","doi":"10.1007/s11030-026-11709-w","DOIUrl":null,"url":null,"abstract":"<p><p>Topoisomerase I (TOP1) is a crucial anticancer target, but the development of traditional TOP1 inhibitors suffers from long research cycles, high costs, and low success rates. Existing artificial intelligence (AI)-driven studies lack systematic comparisons of molecular fingerprints and algorithms, as well as user-friendly predictive application tools. To address these gaps, this study retrieved TOP1 inhibitor activity data from the ChEMBL database, integrated five types of molecular fingerprints (AtomPairs, MACCS, Morgan, PharmacoPFP, and RDKitDes), and constructed and compared classical machine learning (ML) models and deep learning (DL) models, resulting in a total of 40 models. The four top-performing models, SVM::Morgan, RF::Morgan, DNN::MACCS, and KNN::Morgan, achieved ROC-AUC values of 0.93-0.94 under random splitting. Y-scrambling supported that the models learned non-random structure-activity relationships, while SHAP analysis identified key molecular features. The URL of the developed web application is http://drugpred.top:5000 , and this application enables the prediction of TOP1 inhibitory activity via SMILES (Simplified Molecular-Input Line-Entry System) or molecular structure drawing. Additionally, standalone desktop applications (.exe) for offline prediction are freely available at https://github.com/zenghuang8006/TOP1-inhibitor-prediction . Screening of 189,554 SPECS compounds followed by in vitro validation identified AG60 and AI61 as potential TOP1 inhibitors hits, with inhibition rates of 64% and 90% at 400 µM, respectively. Overall, this study provides a practical computational framework for TOP1 inhibitor screening and identifies promising candidate compounds. Notably, scaffold-split AUC values decreased to 0.67-0.82, indicating reduced extrapolative performance for compounds containing previously unseen scaffolds.</p>","PeriodicalId":708,"journal":{"name":"Molecular Diversity","volume":" ","pages":""},"PeriodicalIF":4.3000,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Molecular Diversity","FirstCategoryId":"92","ListUrlMain":"https://doi.org/10.1007/s11030-026-11709-w","RegionNum":2,"RegionCategory":"化学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"CHEMISTRY, APPLIED","Score":null,"Total":0}
引用次数: 0
Abstract
Topoisomerase I (TOP1) is a crucial anticancer target, but the development of traditional TOP1 inhibitors suffers from long research cycles, high costs, and low success rates. Existing artificial intelligence (AI)-driven studies lack systematic comparisons of molecular fingerprints and algorithms, as well as user-friendly predictive application tools. To address these gaps, this study retrieved TOP1 inhibitor activity data from the ChEMBL database, integrated five types of molecular fingerprints (AtomPairs, MACCS, Morgan, PharmacoPFP, and RDKitDes), and constructed and compared classical machine learning (ML) models and deep learning (DL) models, resulting in a total of 40 models. The four top-performing models, SVM::Morgan, RF::Morgan, DNN::MACCS, and KNN::Morgan, achieved ROC-AUC values of 0.93-0.94 under random splitting. Y-scrambling supported that the models learned non-random structure-activity relationships, while SHAP analysis identified key molecular features. The URL of the developed web application is http://drugpred.top:5000 , and this application enables the prediction of TOP1 inhibitory activity via SMILES (Simplified Molecular-Input Line-Entry System) or molecular structure drawing. Additionally, standalone desktop applications (.exe) for offline prediction are freely available at https://github.com/zenghuang8006/TOP1-inhibitor-prediction . Screening of 189,554 SPECS compounds followed by in vitro validation identified AG60 and AI61 as potential TOP1 inhibitors hits, with inhibition rates of 64% and 90% at 400 µM, respectively. Overall, this study provides a practical computational framework for TOP1 inhibitor screening and identifies promising candidate compounds. Notably, scaffold-split AUC values decreased to 0.67-0.82, indicating reduced extrapolative performance for compounds containing previously unseen scaffolds.
期刊介绍:
Molecular Diversity is a new publication forum for the rapid publication of refereed papers dedicated to describing the development, application and theory of molecular diversity and combinatorial chemistry in basic and applied research and drug discovery. The journal publishes both short and full papers, perspectives, news and reviews dealing with all aspects of the generation of molecular diversity, application of diversity for screening against alternative targets of all types (biological, biophysical, technological), analysis of results obtained and their application in various scientific disciplines/approaches including:
combinatorial chemistry and parallel synthesis;
small molecule libraries;
microwave synthesis;
flow synthesis;
fluorous synthesis;
diversity oriented synthesis (DOS);
nanoreactors;
click chemistry;
multiplex technologies;
fragment- and ligand-based design;
structure/function/SAR;
computational chemistry and molecular design;
chemoinformatics;
screening techniques and screening interfaces;
analytical and purification methods;
robotics, automation and miniaturization;
targeted libraries;
display libraries;
peptides and peptoids;
proteins;
oligonucleotides;
carbohydrates;
natural diversity;
new methods of library formulation and deconvolution;
directed evolution, origin of life and recombination;
search techniques, landscapes, random chemistry and more;