Ranran Sun, Jinxin Dong, Hua Jiang, Ruchao Du, Yuxi Zhang
{"title":"CNV-ECOD: A copy number variation detection method based on ECOD algorithm using next-generation sequencing data.","authors":"Ranran Sun, Jinxin Dong, Hua Jiang, Ruchao Du, Yuxi Zhang","doi":"10.1142/S0219720026500083","DOIUrl":"https://doi.org/10.1142/S0219720026500083","url":null,"abstract":"<p><p>Copy number variation (CNV), as a major type of DNA structural variations (SVs), plays a key role in causing human diseases and contributing to genetic diversity. Accurate identification of CNVs is significant for disease mechanism analysis, personalized diagnosis and treatment, and drug development. Although next-generation sequencing (NGS) technology has greatly promoted the development of CNV detection methods, the existing methods generally have problems such as high false positives and inaccurate boundaries. Therefore, a new method is proposed for detecting CNVs in a single sample of NGS data, called CNV-ECOD. The method first employs the empirical-cumulative-distribution-based outlier detection (ECOD) algorithm to identify abnormal signals of read depth (RD) for preliminary detection of CNVs. To correct false positives and refine CNV boundaries further, it integrates paired-end mapping (PEM) and split read (SR) strategies. The integration of the RD-PEM-SR hierarchical progressive framework and the anomaly scoring mechanism based on ECOD can effectively improve the accuracy of CNV detection. Comparing our approach to four peer methods, simulation results demonstrate that it achieves the best balance between precision and sensitivity. Also, the proposed method has the best <i>F</i>1-scores and the highest overlap density scores (ODSs) in real-sample experiments. Therefore, CNV-ECOD is expected to develop into an efficient and robust CNV detection tool.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 3","pages":"2650008"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148296915","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Berkay Orçun Yener, Şurhan Göl, Bora Kutlu, Özgür Öztürk
{"title":"Comparative benchmarking of template-based, evolutionary-diffusion, and generative language models for IsPETase structure prediction.","authors":"Berkay Orçun Yener, Şurhan Göl, Bora Kutlu, Özgür Öztürk","doi":"10.1142/S0219720026510029","DOIUrl":"https://doi.org/10.1142/S0219720026510029","url":null,"abstract":"<p><p>Accurate protein structure prediction is critical for rational enzyme engineering, which requires high-fidelity models. This study benchmarks three distinct structure prediction paradigms against the experimental crystal structure of IsPETase, serving as a diagnostic case study. The evaluated approaches include classical homology modeling (SWISS-MODEL), MSA-conditioned diffusion (AlphaFold 3), and generative language modeling (ESM-3). Predicted models were evaluated using stereochemical validation, molecular docking with a PET dimer analogue, and molecular dynamics simulations. While all approaches reproduced the overall fold and preserved the catalytic triad geometry, notable differences were observed in atomic clashes and hydrogen bonding patterns. ESM-3 showed elevated steric clashes and reduced hydrogen bond counts. Molecular dynamics indicated that the experimental structure maintained the highest stability, with SWISS-MODEL closely following, while ESM-3 displayed greater fluctuations, particularly in loop regions. Crucially, blind docking simulations revealed that the ESM-3 active site was sterically occluded, rendering it inaccessible to the PET dimer. This inaccessibility persisted even after targeted energy minimization. These findings suggest that while generative language models represent a powerful capability for rapid scaffold exploration, they do not yet achieve the thermodynamic precision of established homology and evolutionary approaches required for functional active site engineering.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 3","pages":"2651002"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148296968","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
K Soni Sharmila, S Thanga Revathi, Pokkuluri Kiran Sree
{"title":"Erratum - DDINet: Drug-drug interaction prediction network based on multi-molecular fingerprint features and multi-head attention centered weighted autoencoder.","authors":"K Soni Sharmila, S Thanga Revathi, Pokkuluri Kiran Sree","doi":"10.1142/S0219720026920022","DOIUrl":"10.1142/S0219720026920022","url":null,"abstract":"","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":" ","pages":"2692002"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148200686","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Trap spaces as labelled ideals of SCC posets: A structural-functional theory of reachability in asynchronous boolean networks.","authors":"Belayneh Yibeltal Yizengaw","doi":"10.1142/S0219720026500071","DOIUrl":"https://doi.org/10.1142/S0219720026500071","url":null,"abstract":"<p><p>Boolean networks provide a qualitative framework for modelling regulatory systems when kinetic parameters are unavailable, with cellular phenotypes represented as attractors of the induced dynamics [J. D. Schwab <i>et al.</i>, Concepts in Boolean network modeling, <i>Comput. Struct. Biotechnol. J.</i> <b>18</b>:571-582, 2020]. A central challenge is <i>phenotypic reachability</i>: determining whether asynchronous dynamics can connect invariant regions of the state space, a problem that becomes computationally intractable in large networks [L. Cifuentes-Fontanals, M. Noual, and E. Remy, Revisiting trap spaces in Boolean networks, <i>Theor. Comput. Sci.</i> <b>915</b>:1-20, 2022; K. Perrot and C. Paulevé, Complexity of asynchronous reachability in Boolean networks, <i>Theor. Comput. Sci.</i> <b>1000</b>:114650, 2024.]. We develop a structural theory of reachability in which trap spaces are identified with labelled order ideals of SCC-posets. The SCC-poset determines the order of commitment events, while admissible evaluations encode branching within regulatory modules, so that multistability appears as an intrinsic feature of the theory. Within this framework, we establish necessary and sufficient conditions for reachability, introduce the commitment depth, and show that deciding non-trivial branching is computationally intractable. We further demonstrate that effective interaction structure is jointly determined by topology and Boolean logic. We validate the framework on a Boolean model of CD4[Formula: see text] T-cell differentiation, where refinement chains recover the observed ordering of cytokine response, lineage commitment, and phenotypic branching. In the absence of multistability the structure collapses to a distributive lattice, a non-generic limiting regime.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 3","pages":"2650007"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148297062","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"When pipelines run but coordinates fail: A simple spatial specificity check for false locality in post-GWAS analysis.","authors":"Zheng Han, Hongcheng Zhu, Changzai Li","doi":"10.1142/S0219720026710034","DOIUrl":"https://doi.org/10.1142/S0219720026710034","url":null,"abstract":"<p><p>Some post-GWAS analysis software can run to completion without reporting an error while producing results that are not biologically valid. We call this failure false locality: a result appears to be local to a gene or protein because it was produced from a regional window, but the window or its metadata points to the wrong genomic address. We identify three mechanisms. First, a genome-build address mismatch occurs when GRCh38 protein-QTL coordinates are used with GRCh37 outcome files; in our 91-sentinel audit, 66 coordinates moved by more than 100[Formula: see text]kb and 54 moved by more than 200[Formula: see text]kb after remapping. Second, non-specific regional noise occurs when significant <i>P</i> values persist after the analysis window is deliberately moved to a zero-overlap variant set. Third, location-label blindness occurs when software returns the same output after the declared coordinate label is changed while the SNP table is unchanged; in an official SMR/HEIDI CXCL10 test, the top SNP, SMR <i>P</i> value, and HEIDI <i>P</i> value remained identical across correct, wrong, and shifted labels. We propose a simple Change Test: a result should not be treated as local evidence unless the numerical output changes, weakens, or disappears when the analysis window or coordinate label is intentionally moved to a biologically wrong location. This standard turns software execution from a passive success signal into an explicit spatial-specificity check.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 3","pages":"2671003"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148297026","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Sandip Samaddar, Rituparna Sinha, Rajat Kumar Pal, Rajat Kumar De
{"title":"ReinVar: A model-free paradigm-based reinforcement learning approach to detect copy number variation.","authors":"Sandip Samaddar, Rituparna Sinha, Rajat Kumar Pal, Rajat Kumar De","doi":"10.1142/S021972002650006X","DOIUrl":"10.1142/S021972002650006X","url":null,"abstract":"<p><p>Copy number variation (CNV) is one of the most imperative forms of structural variations that can span over the coding and non-coding regulatory regions of an individual's genome. Copy number variations (CNV) can significantly impact the genotype and phenotype traits by altering the gene dosage, consequently affecting the gene expression landscape concerning various cellular functions and are the cause behind complex diseases in an individual. Exceptionally fast advancement in Next Generation Sequencing (NGS) technology has led to massive growth of DNA-seq data, which contains both Whole Genome Sequence (WGS) and targeted Exome Sequence data of various species including <i>H.sapiens</i>, and precise detection of the DNA region affected by CNV enables the copy number profiling of a genome, thereby understanding our genome. This work has proposed a methodology named ReinVar, which can accurately determine and analyze the underlying copy number profile of the whole genome by adopting a model-free reinforcement learning paradigm. The methodology involves a novel approach to model the problem of identifying CNV as a Markov decision process (MDP), followed by determination of CNV under Reinforcement Learning framework. ReinVar also adopted a Map-Reduce programming paradigm to provide a big data solution to address the issue of exponential growth of NGS read sequence data. ReinVar has shown strong performance in detecting CNV gains and losses across diverse ethnic groups, with a high number of shared variant calls. ReinVar's ability to accurately identify both CNV gains and losses, coupled with consistent detection across ethnic groups and strong ROC characteristics, underscores ReinVar's effectiveness as a robust and sensitive CNV detection method.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 3","pages":"2650006"},"PeriodicalIF":0.8,"publicationDate":"2026-06-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148296901","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"KANBind as a diagnostic probe for DNA-binding protein prediction: A prevalence-calibrated reality check under strict homology control.","authors":"Qipeng Wen, Shaohua Jiang, Yiwen Zhang","doi":"10.1142/S0219720026710022","DOIUrl":"10.1142/S0219720026710022","url":null,"abstract":"<p><p>Deep learning reports over 90% DNA-binding protein (DBP) prediction performance on common benchmarks, but these results are usually obtained on balanced test sets and may not translate to proteome-wide scans with extreme class imbalance. Here, we use KANBind as a diagnostic probe to stress-test sequence-based DBP prediction under strict homology control and realistic prevalence. Evaluated on the homology-controlled HBTD benchmark with prevalence-calibrated reporting, KANBind achieves a calibrated precision of 0.0558 at a realistic bacterial prevalence ([Formula: see text]), implying an expected false discovery rate (FDR) of 94.42%. In a proteome-scale scan, this corresponds to approximately 95 false positives per 100 predicted DBPs. Interpretability analysis indicates that predictions are driven mainly by coarse physicochemical cues such as electrostatics, which may be necessary for DNA binding but are insufficient to determine DBP function. Together, these results suggest that apparent benchmark gains can be dominated by homology leakage and evaluation on balanced sets rather than by generalizable functional rules, motivating stress-test benchmarks with strict homology control and realistic negative backgrounds.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 2","pages":"2671002"},"PeriodicalIF":0.8,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147729862","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Predicting functional co-occurrence probability in PPI networks via multi-level participation expectation.","authors":"Peng Wang","doi":"10.1142/S0219720026500046","DOIUrl":"10.1142/S0219720026500046","url":null,"abstract":"<p><p>This study addresses the problem of protein function annotation and proposes a multi-source biological information-fusion framework called Functional co-Occurrence Probability Estimation (FOPE) for estimating functional co-occurrence probabilities. The framework integrates Protein-Protein Interaction (PPI) network topology and protein domain information, quantifying the functional synergy between protein pairs through bidirectional functional participation modeling. Experiments on four model organisms (A. thaliana, C. elegans, D. melanogaster, and S. cerevisiae) show that FOPE delivers effective predictions for the three main categories of Gene Ontology (GO). Compared to existing representative methods, its macro-Fmax values improved by 25.3%, 19.3%, and 20.9% on average for Biological Process (BP), Cellular Component (CC), and Molecular Function (MF), respectively. Ablation studies further reveal the functional-specific contributions of different information sources: domain information plays a dominant role in MF prediction, while PPI network features are more critical for BP prediction. The effective integration of both is key to achieving comprehensive prediction performance. Robustness tests demonstrate that FOPE maintains strong stability in BP and CC predictions even under significant noise in the PPI network (adding or removing 30% of interactions), verifying the error-tolerance advantages of multi-source information fusion. The FOPE framework proposed in this study provides a feasible information fusion approach for protein function prediction. Preliminary experimental results demonstrate the applicability of the method across different functional categories and multiple model organisms, offering a potential computational pathway for exploring functional synergy relationships among proteins from a system-level perspective.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 2","pages":"2650004"},"PeriodicalIF":0.8,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147730398","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Designing of sorafenib analogs to target c-Raf for the management of hepatocellular carcinoma: Molecular dynamics and mmPBSA analysis.","authors":"Saima Ejaz, Rehan Zafar Paracha, Maryum Nisar, Afreenish Amir, Kashif Saleem, Fouzia Parveen Malik","doi":"10.1142/S0219720025500222","DOIUrl":"10.1142/S0219720025500222","url":null,"abstract":"<p><p><b>Introduction:</b> Sorafenib remains the only approved treatment for advanced hepatocellular carcinoma (HCC), yet its clinical use is hindered by toxicity and the emergence of drug resistance. Sorafenib's anticancer effects are largely attributed to its inhibition of multiple kinases, including c-Raf, a key player in the Ras-Raf-MEK-ERK signaling cascade that promotes cell growth and survival. Given the critical role of c-Raf in tumor progression, targeting this kinase offers a promising strategy for improving therapeutic outcomes. Developing new analogs with stronger c-Raf inhibition, better pharmacokinetics, and reduced side effects could help address the current limitations of sorafenib. <b>Objectives:</b> This study aimed to design novel sorafenib analogs with enhanced binding affinity and favorable pharmacokinetic profiles, specifically targeting the c-Raf kinase to increase therapeutic efficacy against HCC. By using a fragment replacement approach combined with computational methods, the goal was to identify candidates capable of forming stronger, more stable interactions with c-Raf, potentially overcoming resistance linked to sorafenib treatment. <b>Methods:</b> A total of 84 sorafenib analogs (A1-A84) were generated by modifying key functional groups, including the 2-picolinamide and substituted phenyl moieties known to influence kinase binding and anticancer activity. These analogs were evaluated through chemoinformatics and pharmacokinetic screening to assess their drug-likeness and safety. Molecular docking was performed to estimate their binding affinity toward c-Raf. Six top-performing analogs (A2, A6, A9, A20, A22, A63) were selected for further analysis. To evaluate their dynamic behavior, 100[Formula: see text]ns all-atom molecular dynamics simulations were conducted, followed by Molecular Mechanics Poisson-Boltzmann Surface Area (MM-PBSA) calculations to determine binding free energies. Principal component analysis (PCA) was carried out to explore key motion patterns within the protein-ligand complexes. <b>Results:</b> Molecular docking showed that the selected analogs exhibited stronger binding affinities (-11.6 to -10.9[Formula: see text]kcal/mol) compared to sorafenib (-9.3[Formula: see text]kcal/mol) and regorafenib (-9.5[Formula: see text]kcal/mol). Molecular dynamics simulations substantiated the docking results. MM-PBSA results revealed that at 100[Formula: see text]ns, the binding free energy for the c-Raf-sorafenib complex was 86.751[Formula: see text]kJ/mol, while the c-Raf complexes with A2, A6, A9, A20, A22, and A63 demonstrated significantly lower free energies of -129.114, -135.637, -136.242, -127.178, -94.25, and -123.176[Formula: see text]kJ/mol, respectively, indicating stronger and more stable binding. PCA further confirmed the stability and favorable dynamic profiles of these analogs trajectory with c-Raf. <b>Discussion:</b> The improved binding affinities and lower free energies of the top analogs indi","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 2","pages":"2550022"},"PeriodicalIF":0.8,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147729950","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"NNDock2: A neural network-based scoring function for ranking protein-protein docking models.","authors":"Myong-Ho Chae, Gwang So, Ung-Jin Kim","doi":"10.1142/S0219720026500058","DOIUrl":"10.1142/S0219720026500058","url":null,"abstract":"<p><p>Protein-protein interactions (PPIs) play crucial roles in diverse cellular functions and biological processes, and structural knowledge of the protein complexes is valuable for the elucidation of those functions and designing new drugs. Due to the limitations of experimental methods, computational modeling approaches capable of producing reliable protein complex models using molecular docking tools are of considerable practical interest. The success of protein docking largely depends on an accurate scoring function that can pick out good protein docking models. In this work, we present a neural network-based scoring function for scoring protein-protein docking models, NNDock2, the updated version of our previous scoring function, NNDock1. To improve NNDock1, we augmented the training decoys by adding a large number of more distant decoys. In addition, instead of interface root mean square deviation (iRMSD) in NNDock1, we used the fraction of native contact ([Formula: see text] as a target function, which shows better correlation with true model quality. We also applied regularization during training to avoid overfitting. We tested NNDock2 on the protein-protein docking benchmark version 5.0 (BM5), DOCKGROUND dataset, and the CAPRI score set and compared the performance of NNDock2 with other state-of-the-art scoring functions. NNDock2 performed comparably to other state-of-the-art scoring functions, despite the simplicity of the method and low computational costs. We envision that NNDock2 could be used as an independent scoring function or as an element or feature of composite or deep learning-based scoring functions for protein complex model quality estimation.</p>","PeriodicalId":48910,"journal":{"name":"Journal of Bioinformatics and Computational Biology","volume":"24 2","pages":"2650005"},"PeriodicalIF":0.8,"publicationDate":"2026-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"147729993","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":4,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}