Roberto Bernal-Jaquez, Elliot Ridout-Buhl, Emiliano Montoya, León Alday-Toledo, Felipe Aparicio
{"title":"MSI: A Mahalanobis-Based Molecular Similarity Index for High-Dimensional Embeddings.","authors":"Roberto Bernal-Jaquez, Elliot Ridout-Buhl, Emiliano Montoya, León Alday-Toledo, Felipe Aparicio","doi":"10.1002/minf.70051","DOIUrl":"10.1002/minf.70051","url":null,"abstract":"<p><p>Quantifying molecular similarity is crucial for drug discovery and for exploring chemical space. A similarity assessment always combines two independent ingredients: a molecular representation and a similarity coefficient. The most common pairing, binary substructure fingerprints scored with the Tanimoto coefficient, depends strongly on fingerprint bit density and, because it compares unweighted sets of substructure identifiers, is blind to the multiplicity of repeated fragments and frequently returns ranking ties that obscure meaningful chemical relationships. Here, we introduce the Mahalanobis Similarity Index (MSI), which pairs continuous Mol2Vec embeddings with a covariance-aware Mahalanobis distance (d<sub>M</sub>) and an associated Mahalanobis angle (θ<sub>M</sub>) to give a statistically grounded assessment that is invariant under invertible linear reparametrization of the descriptor space. We evaluated MSI on five chemically distinct reference compounds: aspirin, a salicylate nonsteroidal anti-inflammatory drug (NSAID); aniline, an industrial aromatic amine; curcumin, a polyphenolic natural product; ibuprofen, a propionic-acid NSAID; and digitoxin, a cardiac glycoside. Relative to the Tanimoto coefficient computed on ECFP4 fingerprints, MSI improves the analysis in three specific respects: it promotes chemically reasonable analogs that the fingerprint deprioritises; it resolves ranking ties, recovering between 7 and 10 distinct scores among the 10 nearest neighbors where Tanimoto recovers only 3-7; and its geometry varies systematically with HOMO-LUMO energy gaps in the QM9 dataset, indicating that the embedding tracks electronic structure even though it was trained on structural context alone. Polar plots and three-dimensional similarity maps reveal anisotropy within the embedding space and define practical applicability domains for high-similarity retrieval. MSI retains discriminatory power in the regime where the Tanimoto coefficient saturates near zero, and sparse peripheral regions suggest scaffold-hopping opportunities. The dual radial-angular description supports hypothesis-driven reasoning about how structural modifications shift electronic properties. MSI is computationally efficient, chemically interpretable, and offers a practical way to navigate high-dimensional chemical space. Beyond drug discovery, it is applicable to materials science, toxicology, and chemical biology, wherever continuous molecular embeddings are used for property-driven screening.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 9","pages":"e70051"},"PeriodicalIF":3.1,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13522730/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148840830","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Peishuai Xing, Xiaodong Guo, Yang Wang, Zeting Chen, Shicheng Shen, Bin He, Tianyun Li, Dongmei Peng, Zhenhuai Yang
{"title":"Boosting the Prediction Accuracy of Glass Transition Temperature in Polyimides: A Hybrid Machine Learning Approach Integrating Morgan Fingerprints and Molecular Descriptors.","authors":"Peishuai Xing, Xiaodong Guo, Yang Wang, Zeting Chen, Shicheng Shen, Bin He, Tianyun Li, Dongmei Peng, Zhenhuai Yang","doi":"10.1002/minf.70049","DOIUrl":"10.1002/minf.70049","url":null,"abstract":"<p><p>The glass transition temperature (T<sub>g</sub>) of polyimides is a critical parameter determining their processability and application performance. Traditional experimental methods for measuring T<sub>g</sub> are time-consuming and costly, while existing machine learning prediction models predominantly rely on manually defined molecular descriptors, which often fail to fully capture detailed molecular structural information, limiting their prediction accuracy and generalization capability. To address this, this study proposes a hybrid feature engineering strategy combining Morgan fingerprints and molecular descriptors to comprehensively represent the chemical structure of polyimides. Based on a dataset of 1257 polyimide samples from a public database, we systematically compared six feature selection methods and employed multiple mainstream machine learning algorithms for modeling. The results show that the CATB model performed best, achieving a coefficient of determination (R<sup>2</sup>) of 0.882 and a mean absolute error (MAE) of 17.34 °C on an independent test set, with fivefold cross-validation further confirming the model's robustness. SHAP interpretability analysis revealed the significant influence of key features such as the number of rotatable bonds, ether bonds, and ether-linked oxyethylene units on T<sub>g</sub>, providing clear guidance for molecular design. External validation demonstrated the model's strong generalization ability. This study not only achieves high-precision and robust T<sub>g</sub> prediction but also highlights the importance of hybrid feature strategies in polymer property modeling, offering a data-driven foundation for the rational design of polyimides.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70049"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13500066/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148808763","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"PairMap2: A Web Application for Intermediate-Molecule Insertion in Relative Binding Free Energy Calculations.","authors":"Kairi Furui, Masahito Ohue","doi":"10.1002/minf.70046","DOIUrl":"10.1002/minf.70046","url":null,"abstract":"<p><p>PairMap2 is a web application for testing intermediate-molecule insertion in relative binding free energy calculations. Transformations between structurally distant ligands can have insufficient overlap in free-energy space, leading to poor convergence and unreliable predictions. PairMap2 reimplements the PairMap intermediate-insertion workflow as a browser-based tool and supports two-molecule path analysis and multiligand perturbation-map construction. The interface visualizes perturbation networks, generated intermediates, and maximum common substructure-based atom mappings. In nine benchmark cases from the original PairMap study, PairMap2 completed faster in eight cases; for an additional low-similarity transformation (lead optimization mapper (LOMAP) similarity score 0.0111), it achieved an approximately 19.6-fold speedup. PairMap2 enables computational chemists to assess intermediate insertion before free energy perturbation FEB calculations without setting up a local execution environment.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70046"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13454501/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148701873","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Design of Thermotropic Liquid Crystal Molecules With a Wide Temperature Range for the Liquid Crystal Phase Using Machine Learning.","authors":"Naoki Masuyama, Hiromasa Kaneko","doi":"10.1002/minf.70044","DOIUrl":"10.1002/minf.70044","url":null,"abstract":"<p><p>Thermotropic liquid crystals are functional materials whose phase structures change with temperature. They exhibit mesophases-intermediate states between solid and liquid. There is demand for materials that maintain a stable mesophase over a wide temperature range ΔT<sub>LC</sub>, enabling operation even under harsh conditions. Conventional liquid crystal development requires significant time and money from molecular design to physical property evaluation. This study designs liquid crystal molecules exhibiting wide ΔT<sub>LC</sub> in a specific mesophase. The proposed method enables efficient liquid crystal exploration by predicting ΔT<sub>LC</sub> for new materials using machine learning. First, it identifies the mesophase formed from the molecular structure. Next, it predicts the phase transition temperatures (melting point and clearing point) for that phase. Inputting molecules with unknown ΔT<sub>LC</sub> into these models enables the prediction of ΔT<sub>LC</sub> for any given mesophase and efficient exploration of liquid crystal molecules.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70044"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13422677/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148630959","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"MolScope: A Transparent Representation-Aware Workflow Toolkit for Reproducible Medicinal-Chemistry Triage.","authors":"Khaled M Elokely","doi":"10.1002/minf.70048","DOIUrl":"https://doi.org/10.1002/minf.70048","url":null,"abstract":"<p><p>Medicinal-chemistry review is often built from disconnected notebooks, spreadsheets, and slide decks that hide structure-normalization choices and blur the boundary between measured observations and heuristic prioritization. MolScope is an open-source Python workflow toolkit that converts molecular structure collections and optional assay joins into a canonical chemistry table, category-aware summaries, and decision-ready report and picklist artifacts. The workflow makes representation policy explicit, preserves evidence provenance, and packages outputs as portable HTML, Markdown, and CSV bundles. Using frozen example datasets, we show reproducible end-to-end execution, auditable handling of salts, tautomers, stereo ambiguity, charge state, and one bounded round-review example for comparative campaign support.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70048"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13472379/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148761025","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Geometric Deep Learning-Based Drug Design Models for Small-Molecule Drug Discovery.","authors":"Amit Kumar Srivastav, Unnati Modi, Rahul Kumar, Dhiraj Bhatia, Raghu Solanki","doi":"10.1002/minf.70047","DOIUrl":"10.1002/minf.70047","url":null,"abstract":"<p><p>Deep neural network (DNN)-based in silico models show great promise in predicting the properties and bioactivities of novel compounds, including small molecules. Among traditional approaches, structure-based drug design (SBDD) remains a fundamental approach for drug discovery using molecular docking, scoring functions, and molecular dynamics simulations. However, these approaches are often constrained by limited flexibility, resolution, and generalizability. Geometric deep learning (GDL) offers a transformative alternative by enabling models to learn directly from non-Euclidean molecular representations, such as graphs, point clouds, and meshes, capturing critical 3D spatial relationships inherent to protein-ligand interactions. This review highlights the theoretical underpinnings and practical applications of GDL in small-molecule drug discovery, focusing on tasks including binding affinity prediction, virtual screening, de novo molecule generation, pose prediction, ADMET profiling, and protein flexibility modeling. We explore key GDL architectures, graph neural networks, SE(3)-equivariant networks, 3D convolutional neural networks, point cloud models, and geometric transformers, and assess their performance across various drug discovery benchmarks. The integration of geometry-aware AI models with experimental and computational workflows was also highlighted for its potential to streamline hit-to-lead optimization and advance rational drug design. Despite remarkable progress, the field faces challenges including limited high-quality 3D structural datasets, protein flexibility representation, and the interpretability of deep models. Addressing these issues through hybrid modeling approaches, multi-resolution learning, and self-supervised training could further elevate GDL's impact. Ultimately, GDL stands at the frontier of AI-enhanced pharmaceutical innovation, offering unprecedented precision, efficiency, and insight in the pursuit of next-generation therapeutics.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70047"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13454516/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148701813","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Linkerability of Protein Ligands: Insights From Cocrystal Structures and Implications for DNA-Encoded Libraries.","authors":"Raphael M Franzini","doi":"10.1002/minf.70045","DOIUrl":"10.1002/minf.70045","url":null,"abstract":"<p><p>Linkers play a central role in many areas of medicinal chemistry, including proximity inducers, small-molecule conjugates, and DNA-encoded libraries. However, little is known about the accessibility of molecules to linker attachment when bound to proteins. Here, we analyze linker accessibility across protein-ligand complexes in cocrystal structures. A computational workflow was developed to evaluate the linkerability of modifiable positions on molecules based on solvent accessibility, local steric space for introduction of a linker atom, and the geometry of solvent-directed escape paths approximated as conical frustums. Analysis of 8,228 protein-ligand cocrystal structures with 131 431 modifiable positions shows that approximately 22% of positions can accommodate linkers without significant geometric restriction. Limited linkerability of positions influences DEL data and may confound efforts to use such data for lead prediction.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 8","pages":"e70045"},"PeriodicalIF":3.1,"publicationDate":"2026-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13430132/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148664253","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"To What Extent Can We Extrapolate Proteochemometric Models: A Case Study for the SLC6 Transporter Family.","authors":"Uday Abu-Shehab, Gerhard Ecker","doi":"10.1002/minf.70043","DOIUrl":"10.1002/minf.70043","url":null,"abstract":"<p><p>Proteochemometrics (PCM) modeling combines protein and ligand information to create predictive models for biological activity. It aims to extrapolate information across targets, enabling its application in screening drug candidates across a whole family of proteins. In this study, we investigate the ability of PCM models to extrapolate information from data-rich proteins to data-poor ones and present a series of PCM models for the SLC6 transporter family showing reasonable performance (Q2 values up to 0.79). Moreover, feature importance analysis pointed towards residue position A173 in hSERT, corresponding to G149 and G153 in hNET and hDAT, respectively, to be relevant for subtype selectivity. However, examining the impact of different data splits on model validation metrics highlights potential over-optimism when only considering target stratification splits. Target stratification split only maintains the train/test ratio across all targets, preserving per target balance without considering chemical similarity. However, when performing leave-one-transporter-out studies, a considerable drop in performance was observed. This points towards the need for more complex technologies to exploit the potential of PCM and identify new drug candidates.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 7","pages":"e70043"},"PeriodicalIF":3.1,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13413613/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148605323","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"DATTs: A Database of Disease-Associated Therapeutic Targets With Required Actions for Treatment.","authors":"Ryusuke Sawada, Noriko Yuyama Otani, Michio Iwata, Tomokazu Shibata, Shujiro Okuda, Hiroyo Nishide, Ikuo Uchiyama, Yoshihiro Yamanishi","doi":"10.1002/minf.70042","DOIUrl":"10.1002/minf.70042","url":null,"abstract":"<p><p>Proteins involved in pathophysiological mechanisms are widely recognized as promising therapeutic targets for drug discovery. The therapeutic effects of drugs are primarily mediated through the inhibition or activation of target proteins. Therefore, beyond identifying the appropriate molecular target, determining the direction of action on the target protein is a critical factor that shapes the efficacy and characteristics of a treatment. In this paper, we present DATTs (Disease-Associated Therapeutic Targets), which is a new database of therapeutic protein targets with required actions for medical treatment. The information is extracted from public documents such as scientific papers and medical textbooks. The available information includes disease-target protein associations and the actions of drugs on these proteins required for therapeutic effects (e.g., activation, inhibition, or targeting). The DATTs currently includes information on 613 diseases, 1475 therapeutic target proteins, and 7415 disease-protein associations. Additional information such as links to other databases containing data on the molecular functions of target proteins and classification of diseases is also accessible. The DATTs can provide useful clues for investigation of disease mechanisms, novel drug design, and drug repositioning. The DATTs is accessible at https://datts.nibb.ac.jp/.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 7","pages":"e70042"},"PeriodicalIF":3.1,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13413248/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148605365","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Undersampling Techniques for Nonlinear Chemical Space Visualization.","authors":"Akash Surendran, Krisztina Zsigmond, Ramón Alain Miranda-Quintana","doi":"10.1002/minf.70041","DOIUrl":"10.1002/minf.70041","url":null,"abstract":"<p><p>The visualization of high-dimensional chemical space is a critical tool for understanding molecular diversity, structure-property relationships, and for guiding compound selection. However, the performance of non-linear dimensionality reduction (DR) techniques like t-stochastic neighborhood embedding (t-SNE), uniform manifold approximation and projection (UMAP), and generative topographic mapping (GTM) are often susceptible to the choice of hyperparameters, along with the high cost of their training for large datasets. In this study, we investigated the effect of undersampling methods on the choice of hyperparameter selection for these non-linear dimensionality reduction methods. Our results demonstrate that selecting small representative subsets of chemical data not only reduces computational costs associated with hyperparameter training but also serves as an innovative means to train nonlinear DR methods, leading to projections that better preserve the local structure within the chemical space.</p>","PeriodicalId":18853,"journal":{"name":"Molecular Informatics","volume":"45 7","pages":"e70041"},"PeriodicalIF":3.1,"publicationDate":"2026-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13413639/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148605334","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}