{"title":"AResKGLM: a graph-grounded language-model framework for interpretable multi-hop antimicrobial resistance reasoning.","authors":"Jie Ren, Ziyi Yang, Wei Liu, Man Tat Alexander Ng","doi":"10.1093/bib/bbag456","DOIUrl":"10.1093/bib/bbag456","url":null,"abstract":"<p><p>Antimicrobial resistance (AMR) threatens microbiology and microbiome bioinformatics because resistance phenotypes are shaped by interactions among genes, mobile genetic elements, and functional environments across microbial communities. Prioritizing resistance determinants requires models that reason across knowledge graphs (KGs) linking genes, proteins, pathways, drugs, and microbial phenotypes. Existing graph-based methods compress this evidence into scalar scores, whereas large language models can produce explanations not grounded in structured evidence. We developed AResKGLM (Antimicrobial Resistance Knowledge Graph Language Model), a graph-grounded language-model framework for interpretable microbial AMR bioinformatics that serializes breadth-first-search-retrieved multi-hop paths and per-entity biomedical descriptions into a structured Context-Path-Question prompt. Llama-3-8B and DeepSeek-R1-7B are adapted with QLoRA to produce binary link predictions and concise reasoning traces. On the KIDs benchmark, AResKGLM (Llama-3-8B) achieved F1 = 0.8482, outperforming KG-BERT (0.7213), NBFNet (0.5260), and ULTRA (0.2541) (paired Wilcoxon $p = 1.2 times 10^{-7}$). Its advantage increased with reasoning depth: F1 decreased from 0.9197 at 2 hops to 0.8148 at 6 hops, whereas KG-BERT dropped from 0.8110 to 0.6716. Counterfactual path corruption produced an apparent F1 of 0.000, mechanically forced by the probe label assignment; the operative diagnostic is the per-sample flip rate (0.04-0.16), consistent with sensitivity to supplied biological evidence rather than reliance on pretrained priors alone. Cross-species evaluation yielded F1 = 0.81-0.88 with Matthews correlation coefficient (MCC) = 0.35-0.54 on Mycobacterium tuberculosis, Pseudomonas aeruginosa, and Staphylococcus aureus. Temporal ranking of 81 post-2022 gene-drug associations achieved Precision@20 = 100% and AUC-PR = 0.855. AResKGLM offers an interpretable, reproducible framework for multi-hop AMR reasoning, linking candidate prioritization with mechanism-oriented hypothesis generation.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13537466/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148878861","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"X chromosome-wide association studies for quantitative trait loci based on the mixture of general pedigrees and additional unrelated individuals.","authors":"Yi-Fang Wei, Rui-Xiang Zhang, Shun Zhang, Qi Zhong, Yuan-Sheng Li, Jia-Hao Mai, Xian-Bo Wu, Ji-Yuan Zhou","doi":"10.1093/bib/bbag467","DOIUrl":"10.1093/bib/bbag467","url":null,"abstract":"<p><p>Genome-wide association studies have successfully identified many genetic variants associated with complex traits. However, most existing methods target autosomes rather than X chromosome, and several existing X chromosome-wide association studies (XWAS) at quantitative trait loci (QTL) largely focus on unrelated individuals, with limited attention to general pedigrees or mixture of general pedigrees and additional unrelated individuals (called the mixed data for brevity). In this study, we propose nine novel methods for XWAS at QTL in the mixed data (${mathrm{MQX}}_{mathrm{cat}}$, ${mathrm{MQZ}}_{mathrm{max}}$, ${mathrm{MT}}_{mathrm{plinkw}}$, ${mathrm{MT}}_{mathrm{chenw}}$, $mathrm{MwM}3mathrm{VNA}$, ${mathrm{MQMVX}}_{mathrm{cat}}$, ${mathrm{MQMVZ}}_{mathrm{max}}$, $mathrm{MpMV}$, and $mathrm{McMV}$), also applicable to general pedigrees alone. The first four methods test for mean differences across genotypes; the latter four test for differences in both means and variances; $mathrm{MwM}3mathrm{VNA}$ tests for variance differences only. All mean-based and mean-variance-based methods incorporate X chromosome inactivation information, and all nine methods consider genetic relatedness in pedigrees. Simulation studies confirm well-controlled type I error rates, and inclusion of pedigrees significantly improves statistical power. Note that there has been no study focusing on X chromosome for the mixed data or general pedigrees from UK Biobank database, so we apply our proposed methods to this dataset, which identify five total cholesterol (TC)-associated and 13 low-density lipoprotein cholesterol (LDL-C)-associated single nucleotide polymorphisms (SNPs). Linkage disequilibrium (LD) analysis reveals that these SNPs fall into three distinct LD blocks. Functional annotation and gene ontology enrichment analysis reveal 16 and 28 enriched pathways for TC-associated and LDL-C-associated genes, respectively. These methods provide robust and powerful tools for XWAS at QTL in both mixed data and general pedigrees.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13540756/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886290","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Amplification bias in sequencing-based spatial transcriptomics: sources, mechanisms, impacts, and mitigation strategies.","authors":"Yuting Shan, Yanyan Piao, Qinyu Ge","doi":"10.1093/bib/bbag465","DOIUrl":"10.1093/bib/bbag465","url":null,"abstract":"<p><p>Spatial transcriptomics (ST) has emerged as a powerful approach for profiling gene expression in spatial tissue context; yet, its quantitative accuracy remains substantially compromised by amplification bias introduced during the complex library preparation process. These biases arise at multiple stages and accumulate throughout the experimental workflow, distorting transcript abundance, reducing detection sensitivity, and ultimately confounding downstream spatial analyses. This review systematically analyzes amplification bias in ST. We examine how input templates, oligonucleotide components characteristics, enzymatic properties, and experimental conditions collectively contribute to amplification bias, and discuss how these factors propagate through the workflow to generate systematic distortions in data. We further review and critically compare existing strategies for mitigation, encompassing both experimental optimizations and computational approaches and propose a practical decision framework for selecting amplification-bias mitigation strategies according to platform type, sample quality, and RNA input levels. Finally, we outline key challenges and future directions, emphasizing the need for integrative solutions that jointly consider experimental design and computational modeling. This work provides practical guidance for improving data fidelity and interpretation in ST.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 5","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13540725/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148886318","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"G2DR: a genotype-first framework for genetics-informed target prioritization and drug repurposing.","authors":"Muhammad Muneeb, David B Ascher","doi":"10.1093/bib/bbag427","DOIUrl":"10.1093/bib/bbag427","url":null,"abstract":"<p><p>Human genetics offers a scalable route to therapeutic discovery, but practical frameworks that convert genotype-derived signal into ranked target and drug hypotheses remain limited, particularly when matched disease transcriptomics are unavailable. We present G2DR, a genotype-first computational prioritization framework that integrates genetically predicted gene expression, multi-method gene-level testing, pathway enrichment, network context, druggability, and multi-source drug-target evidence to generate hypotheses for downstream follow-up. In a migraine case study of 733 UK Biobank participants (53 cases, 680 controls) using stratified five-fold cross-validation, G2DR imputed genetically regulated expression across seven transcriptome-weight resources and ranked genes using a reproducibility-aware discovery score derived only from training and validation data, followed by a balanced integrated score for target selection. Internal held-out evaluation within the same UK Biobank-derived analytical framework achieved gene-level ROC-AUC of 0.775 and PR-AUC of 0.475 for recovery of test-significant genes, while retaining enrichment for curated migraine-associated biology. Mapping prioritized genes to compounds through Open Targets, DGIdb, and ChEMBL produced drug sets enriched for migraine-linked and literature-associated compounds relative to a global drug background. However, tiered benchmarking showed limited recovery of migraine-specific approved therapies, with stronger signal from mechanism-linked, off-label, and literature-associated pharmacological space. Directionality filtering further distinguished broadly recovered compounds from those with stronger mechanistic compatibility. G2DR is, therefore, best viewed as a modular framework for genetics-informed hypothesis generation in genotype-first settings, not as a clinically actionable target-identification or drug-recommendation system. Prioritized genes and compounds require independent validation.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13455623/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148705539","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Mahmoud Gamal Abdelsalam, Abdulaziz H El-Safty, Amine Zaidi, Tanvir Alam
{"title":"GeneGenie: enhancing biomedical question-answering with agentic graphs.","authors":"Mahmoud Gamal Abdelsalam, Abdulaziz H El-Safty, Amine Zaidi, Tanvir Alam","doi":"10.1093/bib/bbag430","DOIUrl":"10.1093/bib/bbag430","url":null,"abstract":"<p><p>Large language models (LLMs) have revolutionized biomedical research, yet they remain prone to hallucinations and struggle with the precise, multi-hop reasoning required for biomedical analysis. To bridge this gap between generative capability of AI model and factual rigor, this article introduces GeneGenie, a model-agnostic, multi-agent framework built upon a directed acyclic graph architecture. Unlike static prompting strategies, GeneGenie implements a deterministic five-node pipeline that orchestrates query planning, intelligent retrieval-augmented generation across curated databases (GenCC, HGNC, and UniProt), and the dynamic execution of bioinformatics tools, including NCBI E-Utilities and local BLAST+. We evaluated the system using the updated 16-module GeneTuring benchmark, comprising 1600 question-answer pairs. The experimental design compared six state-of-the-art models-including GPT-4o, Claude Sonnet 4.5, and Gemini 2.5 Pro-operating in a standalone \"Direct Mode\" versus the agentic \"Graph Mode.\" The results demonstrate that the graph-based architecture consistently outperforms single-model baselines across all metrics. Notably, among the six selected LLM models we explored, Gemini 2.5 Pro achieved the highest performance, correctly answering 1158 questions (72.375% accuracy), compared with the best baseline score of only 15.8%. Furthermore, our evaluation utilized an \"LLM-as-Judge\" semantic assessment, revealing that the agentic approach significantly enhances not only lexical accuracy but also the completeness and factual grounding of responses. While limitations remain in named entity recognition for protein-coding genes, GeneGenie establishes a robust, reproducible paradigm for future biomedical AI systems, proving that tool-augmented orchestration is superior to reliance on raw model scale alone.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13440129/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148677300","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Prediction of drug hypersensitivity by comprehensive modeling of HLA-peptidomes.","authors":"Yi Zhong, Volker M Lauschke, Yi Wang, Yitian Zhou","doi":"10.1093/bib/bbag350","DOIUrl":"10.1093/bib/bbag350","url":null,"abstract":"<p><p>Human leukocyte antigen (HLA)-B*57:01 associated with abacavir-induced hypersensitivity syndrome (ABC-HSS) is one of the most extensively studied immune-mediated drug hypersensitivity reactions (DHRs). The high odds ratio and strong predictive values of HLA-B*57:01 for ABC-HSS have prompted the Food and Drug Administration and European Medicines Agency to require genetic testing before abacavir treatment. Abacavir binds to HLA-B*57:01 and alters the repertoire of presented peptides, resulting in the activation of autoimmunity. Previous studies employing computational approaches to investigate such DHRs have relied solely on a few crystallized tripartite structures, thus overlooking the full presented peptidome, leading to unsatisfactory predictive results. Here, we employed a state-of-the-art modeling approach to generate HLA structures complexed with over 13 000 presented peptides. We then established a novel computational modeling pipeline to simulate the binding of abacavir to these HLA-peptide complexes. Benchmarking against experimentally determined structures showed that this approach successfully recapitulated the crystalized tripartite structures with high accuracy (RMSD<2.2 Å). We then profiled alterations of the peptide repertoire at key positions in the presence of abacavir and proposed a method that accurately predicts compounds known to trigger T-cell activation. Overall, these results show that comprehensive modeling of the HLA-bound peptidome using advanced structural approaches can enhance the prediction and mechanistic understanding of immune-mediated DHRs.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13331351/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148381662","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"Hierarchical Multi-Omics Trajectory Prediction for fecal microbiota transplantation: a novel machine learning framework for small-sample longitudinal multi-omics integration.","authors":"Zhou Yi-Hui, Sun George","doi":"10.1093/bib/bbag389","DOIUrl":"10.1093/bib/bbag389","url":null,"abstract":"<p><p>Fecal microbiota transplantation (FMT) has emerged as a highly effective treatment for recurrent Clostridioides difficile infection and is being actively investigated for numerous other conditions. While multi-omics studies have revealed dynamic changes in microbial communities and host metabolism following FMT, existing approaches are primarily descriptive and lack the ability to model individual patient trajectories or identify early biomarkers of treatment response. Small-sample, multi-omics, longitudinal prediction presents unique computational challenges: high dimensionality ($p gg n$), multi-omics integration, temporal dynamics, and interpretability. Here, we present Hierarchical Multi-Omics Trajectory Prediction (HMOTP), a purpose-built machine learning framework that addresses these challenges through hierarchical feature construction, multilevel attention mechanisms, and patient-specific trajectory prediction. We evaluated HMOTP on 15 patients with recurrent Clostridioides difficile infection who underwent FMT, with lipidomics and metagenomics profiling at four timepoints spanning 6 months. Notably, naively concatenating multi-omics features degraded Random Forest performance ($93.33%$ to $87.18%$ accuracy), whereas HMOTP's hierarchical integration benefited from the additional omics layer, demonstrating that its advantage stems from structure, not from access to more data. Through hierarchical interpretability, HMOTP identified key biomarkers and revealed cross-omics associations between host lipid metabolism and microbial energy pathways, demonstrating utility for longitudinal modeling and biological discovery in FMT response. HMOTP provides a generalizable, principled framework for personalized medicine applications across small-sample multi-omics problems. Source code and a demo dataset are publicly available.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13380309/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148497030","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"MitoClipSplice: a machine learning framework for resolving mitochondrial RNA cleavage sites from strand-specific RNA-seq soft-clips.","authors":"Qing Yuan, Yu Li, Fanfan Xie, Xinwei Liu, Zhenni Wang, Zhiyang Xu, Yanlin Lin, Gang Wang, Yang Liu, Jinliang Xing, Kaixiang Zhou","doi":"10.1093/bib/bbag429","DOIUrl":"10.1093/bib/bbag429","url":null,"abstract":"<p><p>Mitochondrial RNA processing directed by the transfer ribonucleic acid (tRNA) punctuation model is essential for function and linked to human diseases. Strand-specific RNA sequencing can capture cleavage intermediates as reads with soft-clipping (unmapped sequences at read ends), but these signatures lack systematic characterization, limiting reliable cleavage site identification. We analyzed strand-specific RNA-seq data from 54 samples (35 private, 19 public) encompassing two library types. Soft-clipped reads were evaluated for frequency, quality, guanine-cytosine (GC) content, and fragment size, with sequence-level analysis of clipped portions. We compared random versus non-random priming across 10 sample pairs and assessed alignment strategies. Leveraging multiple features, we developed a random forest model to identify high-confidence cleavage sites and applied it to 20 hepatocellular carcinoma samples. Soft-clipping was prevalent in both library types but significantly higher in second-strand-specific libraries (P < 0.0001), independent of quality metrics. Soft-clipped sequences were predominantly 1-6 nt (87.9%-97.0%), guanine-rich, and preferentially at 3' ends (84.9%-93.8%). Random priming drove high-level 3' soft-clipping on both H-strand (54.47%) and L-strand (28.07%) transcripts, while non-random primers yielded minimal levels (<1.5%). Allowing soft-clipping during alignment increased sequencing depth and precision (P < 0.0001). The random forest model achieved excellent performance (F1 > 0.85, area under the curve > 0.90), with 1-2 nt soft-clips providing the highest signal-to-noise ratio. This first systematic characterization of soft-clipping in mitochondrial RNA-seq establishes a high-fidelity, machine-learning-based workflow for identifying cleavage sites, offering an accessible tool to advance studies of mitochondrial post-transcriptional regulation.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13446511/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148683431","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Yating Li, Xinyue Yu, Hao Zhang, Hao Lin, Bo Liu, Haixia Long
{"title":"AGCLD: an adaptive graph contrastive learning method with denoising for spatial domain identification.","authors":"Yating Li, Xinyue Yu, Hao Zhang, Hao Lin, Bo Liu, Haixia Long","doi":"10.1093/bib/bbag385","DOIUrl":"10.1093/bib/bbag385","url":null,"abstract":"<p><p>Single-cell spatial multi-omics technologies enable the simultaneous acquisition of multimodal molecular profiles and spatial location information in situ, providing a novel perspective for spatial domain identification and functional characterization of tissues. However, existing methods still suffer from several limitations, including insufficient denoising capability for single-cell data, reliance on static graph structures, and inadequate exploitation of the complementary relationships between spatial information and molecular features. To address these challenges, we propose AGCLD, an adaptive graph contrastive learning method with denoising for spatial domain identification. Specifically, a modality-specific denoising variational autoencoder is first employed to learn robust latent representations, thereby effectively mitigating noise interference. A differentiable graph generator is then introduced to adaptively construct spatial adjacency graphs and expression similarity graphs, alleviating the bias introduced by fixed neighborhood assumptions. Finally, AGCLD utilizes a dual-graph attention network to encode the spatial adjacency and expression similarity graphs, yielding spatial and feature embeddings, and incorporates a contrastive learning mechanism to align the dual-view representations, thereby enhancing representation consistency. Extensive experiments on five spatial multi-omics datasets demonstrate that AGCLD outperforms state-of-the-art methods, including SpatialGlue, in spatial domain identification tasks.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13367445/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148444505","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
{"title":"DeepACPred: an integrated multistage framework for anticancer peptide discovery and activity prediction.","authors":"Bo Zhang, Ruifang Li, Kedong Yin, Yufeng Yang, Jinhua Zhang, Mengwan Jiang, Huijie Wang, Shiyu Li, Lujing Jia","doi":"10.1093/bib/bbag424","DOIUrl":"10.1093/bib/bbag424","url":null,"abstract":"<p><p>Artificial intelligence accelerates anticancer peptides (ACPs) discovery. However, existing computational methods lack integration of identification with activity-based candidate prioritization. Here, we present DeepACPred, a three-stage pipeline encompassing ACP binary classification model, ACP multilabel classification model, and ACP IC50 prediction model, leveraging multimodal features from ESM2 protein language model embeddings, AAindex physicochemical descriptors, and sequence composition. On 5712 benchmark sequences, the binary classifier achieved 95.10% accuracy (AUC = 0.9913), with performance remaining stable under CD-HIT cluster-aware splitting at 40%-90% identity thresholds. Multilabel cancer-type prediction yielded macro-F1 = 0.9124 across seven cancer types, and log10(IC50) regression achieved Spearman ρ = 0.8602 under 5-fold cross-validation. Ablation experiments showed task-dependent feature contributions rather than uniformly additive multimodal effects. Applied to 260 000 motif-enriched 18-mer candidates, DeepACPred selected 12 peptides predicted to be active against breast cancer cells, all of which showed measurable in vitro cytotoxic activity against murine 4T1 cells in OD-derived dose-response assays (IC50: 0.88-36.83 μg/ml). Although prospective IC50 ranking showed limited fine-grained resolution, these results support the use of the regression module for coarse candidate enrichment. In conclusion, DeepACPred provides a systematic framework for ACP candidate enrichment and prioritization.</p>","PeriodicalId":9209,"journal":{"name":"Briefings in bioinformatics","volume":"27 4","pages":""},"PeriodicalIF":7.3,"publicationDate":"2026-07-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13452478/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148697042","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}