BMC Bioinformatics最新文献

筛选
英文 中文
AssayBLAST v2: major update improving reliability and reporting of the in silico analysis of molecular multi-parameter assays. AssayBLAST v2:主要更新提高了分子多参数分析的可靠性和报告。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-31 DOI: 10.1186/s12859-026-06595-w
Tom Eulenfeld, Maximillian Collatz, Sascha D Braun, Ralf Ehricht
{"title":"AssayBLAST v2: major update improving reliability and reporting of the in silico analysis of molecular multi-parameter assays.","authors":"Tom Eulenfeld, Maximillian Collatz, Sascha D Braun, Ralf Ehricht","doi":"10.1186/s12859-026-06595-w","DOIUrl":"10.1186/s12859-026-06595-w","url":null,"abstract":"<p><strong>Introduction: </strong>Accurate in silico evaluation of primers and probes is essential for the rational design of molecular multi-parameter assays. We present AssayBLAST v2 to automate and simplify this process for extensive assay designs.</p><p><strong>Results: </strong>A newly integrated strand and proximity check enables precise validation of corresponding oligonucleotides, ensuring correct orientation and spacing required for amplification. Based on predicted oligonucleotide interactions, AssayBLAST v2 determines the theoretical amplification outcomes, offering a computational benchmark for downstream wet-lab validation and performance correlation. Additionally, the updated software integrates an adaptive BLAST parameter optimization that dynamically scales with database size, thereby improving both analytical sensitivity and computational performance. These improvements are supported by a comparative evaluation against the previous version of AssayBLAST.</p><p><strong>Conclusions: </strong>Collectively, these enhancements streamline the assay development workflow, reduce costs associated with suboptimal primer and probe synthesis, and increase the robustness and reliability of molecular diagnostics and research applications.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13531902/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148872977","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Correction: Semi-automatic geometrical reconstruction and analysis of filopodia dynamics in 4D two photon microscopy images. 校正:四维双光子显微镜图像中丝足动力学的半自动几何重建和分析。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-25 DOI: 10.1186/s12859-026-06557-2
Blaž Brence, Josephine Brummer, Vincent J Dercksen, Mehmet Neset Özel, Abhishek Kulkarni, Neele Wolterhoff, Steffen Prohaska, Peter Robin Hiesinger, Daniel Baum
{"title":"Correction: Semi-automatic geometrical reconstruction and analysis of filopodia dynamics in 4D two photon microscopy images.","authors":"Blaž Brence, Josephine Brummer, Vincent J Dercksen, Mehmet Neset Özel, Abhishek Kulkarni, Neele Wolterhoff, Steffen Prohaska, Peter Robin Hiesinger, Daniel Baum","doi":"10.1186/s12859-026-06557-2","DOIUrl":"10.1186/s12859-026-06557-2","url":null,"abstract":"","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13505034/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148817155","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
TOMOC: a topology-driven multi-objective evolutionary framework for robust overlapping protein complex detection in noisy PPI networks. TOMOC:一个拓扑驱动的多目标进化框架,用于噪声PPI网络中鲁棒重叠蛋白复合物检测。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-18 DOI: 10.1186/s12859-026-06603-z
Mustafa Abbas, David Broneske, Gunter Saake
{"title":"TOMOC: a topology-driven multi-objective evolutionary framework for robust overlapping protein complex detection in noisy PPI networks.","authors":"Mustafa Abbas, David Broneske, Gunter Saake","doi":"10.1186/s12859-026-06603-z","DOIUrl":"https://doi.org/10.1186/s12859-026-06603-z","url":null,"abstract":"<p><strong>Background: </strong>Protein complexes constitute fundamental functional modules within cells and play a crucial role in regulating many biological processes. Detecting protein complexes from protein-protein interaction (PPI) networks has therefore become a central problem in computational systems biology. However, many existing computational approaches struggle to accurately identify overlapping complexes where proteins participate in multiple functional modules simultaneously. In addition, large-scale PPI networks are inherently noisy and incomplete due to experimental limitations, which significantly affects the reliability of complex detection methods.</p><p><strong>Results: </strong>In this work, we propose TOMOC, a topology-driven multi-objective evolutionary framework for robust detection of overlapping protein complexes in noisy PPI networks. The novelty of TOMOC lies in the integration of a topology-driven bi-objective formulation, an edge-based evolutionary representation that naturally supports overlapping memberships, and a topology-aware structural refinement mechanism within a unified framework for protein complex detection in noisy PPI networks. The proposed framework introduces an edge-based evolutionary representation that models candidate solutions at the interaction level, allowing overlapping memberships to emerge naturally during decoding. It further optimizes two complementary structural objectives by minimizing average conductance and triangle-density loss, enabling the algorithm to balance boundary quality and internal structural density. In addition, a topology-aware structural overlap refinement (SOR) operator is designed to improve structural coherence and robustness against noisy interactions through boundary-aware repair, triangle-closure expansion, and triangle-support pruning. Extensive experiments conducted on three benchmark PPI networks (Yeast-D1, Yeast-D2, and Collins) demonstrate that TOMOC achieves competitive performance compared with several state-of-the-art methods in terms of precision, recall, and F1-score.</p><p><strong>Conclusions: </strong>The proposed TOMOC framework provides an effective and scalable topology-driven approach for detecting overlapping protein complexes directly from PPI network topology. By integrating multi-objective evolutionary optimization with topology-aware refinement mechanisms, TOMOC effectively captures the structural characteristics of protein complexes and demonstrates strong robustness when applied to large and noisy biological interaction networks.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13488024/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148787890","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
MiRQuery: a user-friendly web app for the interactive analysis and visualization of microRNA sequencing data. MiRQuery:一个用户友好的web应用程序,用于microRNA测序数据的交互式分析和可视化。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-17 DOI: 10.1186/s12859-026-06479-z
Julianne C Yang, Jake Sauter, Gregory C Adam, Richard Carr
{"title":"MiRQuery: a user-friendly web app for the interactive analysis and visualization of microRNA sequencing data.","authors":"Julianne C Yang, Jake Sauter, Gregory C Adam, Richard Carr","doi":"10.1186/s12859-026-06479-z","DOIUrl":"https://doi.org/10.1186/s12859-026-06479-z","url":null,"abstract":"<p><strong>Background: </strong>MicroRNAs (miRNAs) are a class of small noncoding RNAs that inhibit the translation of target messenger RNAs (mRNAs). Given that a single miRNA can regulate the translation of many mRNAs, miRNAs have emerged as critical regulators of physiological processes. MiRNAs have been linked to the development and progression of cancers, neurodegenerative and other diseases, most recently using high-throughput miRNA \"miRNome\" sequencing. As miRNome sequencing represents a newer 'omics application, limited guidance is available for how to analyze this data. Existing interfaces that enable non-computational users to interpret and perform comprehensive secondary analysis on their own miRNome data are limited in functionality and/or interactivity. Therefore, we developed MiRQuery to address this need.</p><p><strong>Results: </strong>MiRQuery is an RShiny application which features common visualization methods for high-throughput sequencing data, such as multidimensional scaling, stacked column charts, heatmaps, and boxplots to compare expression across groups for a user-specified miRNA of interest. MiRQuery further provides support for differential miRNA and gene expression analysis. Unique to miRNome sequencing data analysis, users may retrieve predicted gene targets of differentially expressed miRNA and follow up with pathway overrepresentation analysis of the gene targets. Finally, if users upload paired bulk mRNA sequencing data, they may identify differentially expressed genes and negatively correlated miRNA-gene pairs.</p><p><strong>Conclusions: </strong>By providing access to sophisticated bioinformatics tools through a user-friendly interface, MiRQuery empowers both scientists new to bioinformatics and bioinformaticians new to the field to extract insights rapidly and reproducibly from their sequencing data. MiRQuery can be accessed through PositConnect at https://julianneyang-mirquery.share.connect.posit.cloud/ , and alternatively is available by user local installation via instructions on the Github project homepage.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13483514/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148787929","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
plsMD: a plasmid reconstruction tool from short-read assemblies. plsMD:一个质粒重建工具,从短读片段。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-13 DOI: 10.1186/s12859-026-06585-y
Maryam Lotfi, Deena Jalal, Ahmed A Sayed
{"title":"plsMD: a plasmid reconstruction tool from short-read assemblies.","authors":"Maryam Lotfi, Deena Jalal, Ahmed A Sayed","doi":"10.1186/s12859-026-06585-y","DOIUrl":"https://doi.org/10.1186/s12859-026-06585-y","url":null,"abstract":"<p><strong>Background: </strong>While whole genome sequencing has become a cornerstone of antimicrobial resistance surveillance, the reconstruction of plasmid sequences from short-read data remains a challenge due to repetitive sequences and assembly fragmentation. Current computational tools for plasmid identification and binning have limitations in reconstructing full plasmid sequences, hindering downstream analyses like phylogenetic studies and antimicrobial resistance gene tracking.</p><p><strong>Results: </strong>We present plsMD, a tool designed for full plasmid reconstruction from short-read assemblies. plsMD integrates Unicycler assemblies with replicon and full plasmid sequence databases to guide plasmid reconstruction through a series of contig manipulations. Using two datasets - an established benchmark dataset used in previous benchmarking studies and a novel dataset consisting of newly sequenced bacterial isolates - plsMD outperformed existing tools in both. In the benchmark dataset, it achieved excellent recall, precision, and F1 scores of 91.3%, 95.5%, and 92.0%, respectively. In the novel dataset, it achieved recall, precision, and F1 scores of 77.6, 88.9 and 74.5%, respectively. plsMD supports two usage modalities: single-sample analysis for plasmid reconstruction and gene annotation, and batch-sample analysis for phylogenetic investigations of plasmid transmission.</p><p><strong>Conclusions: </strong>plsMD represents a significant advancement in plasmid analysis, offering a robust solution for utilizing existing short-read whole genome sequencing data to study plasmid-mediated antimicrobial resistance spread and evolution.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13474732/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148765878","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Introducing GalaxyR: an easy-to-use R implementation of the Galaxy API. 介绍GalaxyR:一个易于使用的Galaxy API的R实现。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-10 DOI: 10.1186/s12859-026-06591-0
Julian Frey, Zoe Schindler, Muhammad Ali, Joshua Braun-Wimmer, Kilian Gerberding, Katja Kröner, Elena Larysch, Daniel Lusk, Janusch Vanja-Jehle, Yannik Wardius, Maximilian Weidenfeller, Teja Kattenborn, Thomas Seifert
{"title":"Introducing GalaxyR: an easy-to-use R implementation of the Galaxy API.","authors":"Julian Frey, Zoe Schindler, Muhammad Ali, Joshua Braun-Wimmer, Kilian Gerberding, Katja Kröner, Elena Larysch, Daniel Lusk, Janusch Vanja-Jehle, Yannik Wardius, Maximilian Weidenfeller, Teja Kattenborn, Thomas Seifert","doi":"10.1186/s12859-026-06591-0","DOIUrl":"10.1186/s12859-026-06591-0","url":null,"abstract":"<p><strong>Background: </strong>Standardisation, accessibility, and reproducibility remain persistent challenges in data‑intensive scientific research. The Galaxy platform addresses these issues by providing a web-based environment for executing, sharing, and publishing computational workflows, yet programmatic access has been largely centred on the Python ecosystem through BioBlend.</p><p><strong>Methods: </strong>Here, we introduce GalaxyR, a native R package that provides a comprehensive and structured interface to the Galaxy application programming interface. GalaxyR enables users to manage histories, upload and retrieve data, execute tools and workflows, monitor jobs, and inspect results directly from within the R environment. By integrating Galaxy's scalable computational infrastructure with R's widely adopted data analysis ecosystem, GalaxyR facilitates automated, reproducible, and resource‑efficient workflows without requiring local high‑performance computing resources. The package supports both HTTPS‑ and FTP‑based data transfer, robust history and dataset management, and programmatic workflow orchestration across Galaxy instances.</p><p><strong>Results: </strong>An application example demonstrates the use of GalaxyR for large-scale processing of drone-based laser scanning data, where computationally intensive tree‑level segmentation was delegated to the Galaxy infrastructure, while workflow control and preprocessing were handled in R within a few lines of code.</p><p><strong>Conclusions: </strong>GalaxyR thus bridges a critical gap for R users, significantly expanding access to Galaxy‑based analyses and enabling scalable, reproducible research across bioinformatics, ecology, remote sensing, and related data-driven disciplines.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13465283/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148711316","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Machine learning-based prediction of SARS-CoV-2 bioactivity: integrating IC50 regression and activity classification using multi-task neural networks. 基于机器学习的SARS-CoV-2生物活性预测:基于多任务神经网络集成IC50回归和活性分类
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-08-05 DOI: 10.1186/s12859-026-06573-2
Aya I Maiyza, Sohila Osama, Hanan A Hassan
{"title":"Machine learning-based prediction of SARS-CoV-2 bioactivity: integrating IC<sub>50</sub> regression and activity classification using multi-task neural networks.","authors":"Aya I Maiyza, Sohila Osama, Hanan A Hassan","doi":"10.1186/s12859-026-06573-2","DOIUrl":"10.1186/s12859-026-06573-2","url":null,"abstract":"<p><p>Accurate prediction of compound bioactivity is essential for accelerating antiviral drug discovery and reducing experimental costs. Machine learning (ML) methods have shown considerable promise in modeling structure-activity relationships and compound potency. In this study, we present an integrated ML framework for predicting IC<sub>50</sub> and pIC<sub>50</sub> values of compounds active against SARS-CoV-2, key indicators of antiviral potency. The proposed framework comprises three complementary approaches: (i) a regression model for quantitative IC<sub>50</sub> prediction validated against experimental data; (ii) a classification model that categorizes compounds into active and inactive classes to support compound prioritization; and (iii) a multi-task neural network that jointly performs IC<sub>50</sub> regression and activity classification, enhancing predictive performance and interpretability. A distinctive feature of this work is the incorporation of ligand efficiency (LE) as a criterion for activity classification, offering an alternative perspective on compound prioritization that has not been previously explored in SARS-CoV-2 bioactivity modeling. The proposed models demonstrate strong predictive capability, achieving a coefficient of determination ([Formula: see text]) of 0.77 using a neural network with feature selection, while the Random Forest classifier attains an accuracy, precision, and recall of approximately 0.92. These results highlight the potential of integrated regression, classification, and multi-task learning approaches as scalable and cost-effective tools for SARS-CoV-2 bioactivity prediction and antiviral drug discovery.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13445897/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148677037","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Residual-stream geometry of single-cell foundation models carries incremental gene-regulatory signal across tissues. 单细胞基础模型的残流几何结构携带跨组织的增量基因调控信号。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-07-24 DOI: 10.1186/s12859-026-06538-5
Ihor Kendiukhov
{"title":"Residual-stream geometry of single-cell foundation models carries incremental gene-regulatory signal across tissues.","authors":"Ihor Kendiukhov","doi":"10.1186/s12859-026-06538-5","DOIUrl":"10.1186/s12859-026-06538-5","url":null,"abstract":"<p><strong>Background: </strong>Single-cell foundation models such as scGPT and Geneformer learn rich representations of gene expression programs, but whether these representations encode gene regulatory relationships beyond expression-level confounds remains unclear. Attention patterns in these models have been shown to capture co-expression rather than direct regulation, leaving open the question of whether deeper representations-particularly the residual stream-contain genuine regulatory information.</p><p><strong>Results: </strong>We systematically investigated residual-stream geometry in scGPT and Geneformer across four tissue contexts from the Tabula Sapiens atlas, evaluating whether geometric proximity between gene vectors provides incremental predictive value for curated TRRUST transcription factor-target edges beyond expression confounds. Under repeated stratified cross-validation, geometric features provided significant incremental signal in kidney and immune settings, validated by label-permutation and geometry-shuffle null controls; centered-cosine similarity, PCA projection and multi-layer bundling recovered comparable signal in lung tissues, and the multi-layer bundle improved every domain (kidney ΔAUROC = + 0.122, immune + 0.042, lung + 0.028, external lung + 0.027; geometry-augmented AUROC 0.60-0.69). The effect was fully robust to leave-TF-out and leave-target-out cross-validation and to harder degree- and expression-matched negative edges, but under the stricter leave-both-out split-no transcription factor and no target shared between folds-it collapsed to near-zero (ΔAUROC at most + 0.003, and not statistically significant in kidney or immune), marking the ceiling of out-of-entity generalization. With a comparable per-layer residual-stream extraction applied to both models, the apparent Geneformer advantage mostly disappeared (small residual gaps remained in three of four domains), indicating it largely reflected representation-construction choices rather than a substantial architectural difference. Asymmetric geometric features predicted regulatory edge orientation (AUROC 0.80-0.90), and the geometric signal added incremental value on top of expression-based gene regulatory network (GRN) inference (GENIE3, co-expression).</p><p><strong>Conclusion: </strong>Foundation model residual streams carry incremental, regulatory-relevant geometric signal that is distributed across layers and that complements expression-based GRN inference for retrospective edge prioritization. The signal is statistical enrichment rather than a stand-alone regulatory classifier: absolute performance is modest and out-of-entity generalization is limited, so its practical role is as an orthogonal evidence channel for edge re-ranking and hypothesis prioritization in multi-evidence frameworks.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-07-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13418759/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148618318","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Mol2Image: an enhanced DDI prediction framework leveraging drug molecular descriptors. Mol2Image:利用药物分子描述符的增强型DDI预测框架。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-07-22 DOI: 10.1186/s12859-026-06552-7
Nourhan Helmy, Huda Amin Maghawry, Nagwa Badr
{"title":"Mol2Image: an enhanced DDI prediction framework leveraging drug molecular descriptors.","authors":"Nourhan Helmy, Huda Amin Maghawry, Nagwa Badr","doi":"10.1186/s12859-026-06552-7","DOIUrl":"10.1186/s12859-026-06552-7","url":null,"abstract":"<p><p>Drug-drug interactions (DDIs) are a critical safety issue in clinical practice, as they can lead to severe and often unpredictable adverse effects. This risk becomes significantly higher in multi-drug therapies, which are increasingly used in the treatment of complex and chronic diseases such as cancer, cardiovascular disorders, and diabetes. However, identifying DDIs through in vivo studies is costly and time-consuming. In this study, a novel DDI prediction model, Mol2Image, has been proposed that utilizes chemical structure features derived from Simplified Molecular Input Line Entry System (SMILES) representations, including molecular property descriptors and structural fingerprints. The proposed model combines chemical structure information with automated feature learning. Molecular descriptors and structural fingerprints extracted from SMILES representations are converted into visual patterns that capture key chemical characteristics of each drug. These images are then processed by a Convolutional Neural Network (CNN) to learn high-level structural features associated with drug-drug interactions. The model is trained and evaluated using two benchmark DDI datasets: the Drugbank dataset, which consists of 443,046 interactions, and ChCh-Miner, which consists of 48,514 DDIs. Experimental results demonstrate that the proposed model (Mol2Image) achieves competitive performance compared with several state-of-the-art methods. Experimental results demonstrate that the proposed model consistently outperforms existing approaches, achieving accuracies of 0.9608 and 0.9683 using the Drugbank dataset and ChCh-Miner dataset, respectively. Ultimately, Mol2Image provides a highly scalable, strictly structure-centric framework that ensures superior predictive accuracy with minimal computational overhead, operating entirely independently of clinical data.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-07-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13393720/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148576944","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2. 利用DNABERT和DNABERT-2对酿酒酵母DNA复制起源的可解释性预测。
IF 4.4 3区 生物学
BMC Bioinformatics Pub Date : 2026-07-22 DOI: 10.1186/s12859-026-06562-5
Zohreh Piroozeh, Ildem Akerman, Olga V Kalinina, Stefan Kesselheim, Alina Bazarova
{"title":"Interpretable prediction of DNA replication origins in S. cerevisiae using DNABERT and DNABERT-2.","authors":"Zohreh Piroozeh, Ildem Akerman, Olga V Kalinina, Stefan Kesselheim, Alina Bazarova","doi":"10.1186/s12859-026-06562-5","DOIUrl":"10.1186/s12859-026-06562-5","url":null,"abstract":"<p><strong>Background: </strong>DNA replication is a biological process in which a single DNA molecule is duplicated, initiating from multiple genomic sites known as replication origins. Identifying replication origins and analyzing their underlying base sequence composition is crucial for understanding the mechanisms of DNA replication. Although there are various machine learning and deep learning approaches for origin prediction, many rely on labor intensive feature engineering or lack interpretability. We fine-tune two genome-based pretrained language models, DNABERT and DNABERT-2, to predict replication origins in budding yeast and unravel the DNA base composition behind them. The key contribution of this study is a systematic framework for analyzing genomic language models for replication origin prediction, combining controlled dataset design with model-specific explainability pipelines to examine how different tokenization strategies influence learned sequence features and whether such approaches can highlight biologically meaningful signals.</p><p><strong>Results: </strong>We evaluate both models on the designed datasets to ensure robustness and support explainability. DNABERT demonstrates consistent performance, achieving an average accuracy of 0.72 for more challenging and 0.83 for the easier dataset. In comparison, DNABERT-2 achieved comparable scores of 0.72 and 0.81 on the same datasets. Our attention-based motif discovery pipeline enhances the interpretability of DNABERT, by identifying motifs from high-attention fragments that closely match known sequence patterns of replication origins. Perturbation-based explanation methods, including Shapley additive explanations, were applied to interpret DNABERT-2's learning mechanism. This analysis identified tokens with high attribution scores aligned with biologically relevant sequence composition.</p><p><strong>Conclusion: </strong>Our study demonstrates that both models identify replication origin sequences, albeit through different learning strategies. Tokenization appears to influence model learning and attention behavior in these models. The overlapping k-mer tokenization used in DNABERT yields more interpretable attention maps compared to the byte pair encoding tokenization employed in DNABERT-2. We show that despite sharing the same BERT-style architecture, DNABERT captures relevant short-range patterns and some sequence dependencies beyond just local context, as reflected in its attention maps. In contrast, DNABERT-2's alternative tokenization strategy biases its learning toward relevant short-range patterns by optimizing token weighting.</p>","PeriodicalId":8958,"journal":{"name":"BMC Bioinformatics","volume":"27 1","pages":""},"PeriodicalIF":4.4,"publicationDate":"2026-07-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13393499/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148560708","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":3,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
相关产品
×
本文献相关产品
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书