{"title":"基于深度标签相关性和标签歧义的文本分类多标签特征选择","authors":"Gurudatta Verma, Tirath Prasad Sahu","doi":"10.1016/j.engappai.2025.110403","DOIUrl":null,"url":null,"abstract":"<div><div>Multi-label text classification, where each document can be associated with multiple labels simultaneously, poses unique challenges in feature selection due to the complex relationships between features and labels. In this paper, we propose a novel <strong>D</strong>eep Label <strong>R</strong>elevance and <strong>L</strong>abel <strong>A</strong>mbiguity (DLRLA) based multi-label feature selection method designed for multi-label text data. Our approach constructs a quasi-relevance matrix integrating low-order, high-order feature-label relevance and label ambiguity. The low-order relevance captures the direct association between individual features and labels, while the high-order relevance accounts for the interactions between feature combinations and labels, collectively termed as deep label relevance. Label ambiguity, measured using information entropy, quantifies the uncertainty associated with each label. The quasi-relevance matrix is then evaluated using Grey Relation Optimization to rank and select the most informative features based on multiple relevance criteria. Additionally, feature-feature relevance is incorporated to reduce the candidate set of high-order features, mitigating computational complexity. Elastic Net Regression, a linear regularized model, estimates feature-label relevance, enabling efficient feature selection while addressing multicollinearity. For multi-label classification, we leverage the Multi-Label K-Nearest Neighbors algorithm, where the key parameters (number of neighbours k and smoothing factor s) are optimized using Particle Swarm Optimization. The proposed DLRLA method is extensively evaluated on ten multi-label text benchmark datasets, considering six performance evaluation metrics. Comparative analyses with seven state-of-the-art methods are conducted. Furthermore, a stability analysis of DLRLA is performed across all datasets and evaluation metrics, showcasing its robustness and consistency.</div></div>","PeriodicalId":50523,"journal":{"name":"Engineering Applications of Artificial Intelligence","volume":"148 ","pages":"Article 110403"},"PeriodicalIF":8.0000,"publicationDate":"2025-03-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Deep label relevance and label ambiguity based multi-label feature selection for text classification\",\"authors\":\"Gurudatta Verma, Tirath Prasad Sahu\",\"doi\":\"10.1016/j.engappai.2025.110403\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><div>Multi-label text classification, where each document can be associated with multiple labels simultaneously, poses unique challenges in feature selection due to the complex relationships between features and labels. In this paper, we propose a novel <strong>D</strong>eep Label <strong>R</strong>elevance and <strong>L</strong>abel <strong>A</strong>mbiguity (DLRLA) based multi-label feature selection method designed for multi-label text data. Our approach constructs a quasi-relevance matrix integrating low-order, high-order feature-label relevance and label ambiguity. The low-order relevance captures the direct association between individual features and labels, while the high-order relevance accounts for the interactions between feature combinations and labels, collectively termed as deep label relevance. Label ambiguity, measured using information entropy, quantifies the uncertainty associated with each label. The quasi-relevance matrix is then evaluated using Grey Relation Optimization to rank and select the most informative features based on multiple relevance criteria. Additionally, feature-feature relevance is incorporated to reduce the candidate set of high-order features, mitigating computational complexity. Elastic Net Regression, a linear regularized model, estimates feature-label relevance, enabling efficient feature selection while addressing multicollinearity. For multi-label classification, we leverage the Multi-Label K-Nearest Neighbors algorithm, where the key parameters (number of neighbours k and smoothing factor s) are optimized using Particle Swarm Optimization. The proposed DLRLA method is extensively evaluated on ten multi-label text benchmark datasets, considering six performance evaluation metrics. Comparative analyses with seven state-of-the-art methods are conducted. Furthermore, a stability analysis of DLRLA is performed across all datasets and evaluation metrics, showcasing its robustness and consistency.</div></div>\",\"PeriodicalId\":50523,\"journal\":{\"name\":\"Engineering Applications of Artificial Intelligence\",\"volume\":\"148 \",\"pages\":\"Article 110403\"},\"PeriodicalIF\":8.0000,\"publicationDate\":\"2025-03-08\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Engineering Applications of Artificial Intelligence\",\"FirstCategoryId\":\"94\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S0952197625004038\",\"RegionNum\":2,\"RegionCategory\":\"计算机科学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"AUTOMATION & CONTROL SYSTEMS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Engineering Applications of Artificial Intelligence","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0952197625004038","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"AUTOMATION & CONTROL SYSTEMS","Score":null,"Total":0}
Deep label relevance and label ambiguity based multi-label feature selection for text classification
Multi-label text classification, where each document can be associated with multiple labels simultaneously, poses unique challenges in feature selection due to the complex relationships between features and labels. In this paper, we propose a novel Deep Label Relevance and Label Ambiguity (DLRLA) based multi-label feature selection method designed for multi-label text data. Our approach constructs a quasi-relevance matrix integrating low-order, high-order feature-label relevance and label ambiguity. The low-order relevance captures the direct association between individual features and labels, while the high-order relevance accounts for the interactions between feature combinations and labels, collectively termed as deep label relevance. Label ambiguity, measured using information entropy, quantifies the uncertainty associated with each label. The quasi-relevance matrix is then evaluated using Grey Relation Optimization to rank and select the most informative features based on multiple relevance criteria. Additionally, feature-feature relevance is incorporated to reduce the candidate set of high-order features, mitigating computational complexity. Elastic Net Regression, a linear regularized model, estimates feature-label relevance, enabling efficient feature selection while addressing multicollinearity. For multi-label classification, we leverage the Multi-Label K-Nearest Neighbors algorithm, where the key parameters (number of neighbours k and smoothing factor s) are optimized using Particle Swarm Optimization. The proposed DLRLA method is extensively evaluated on ten multi-label text benchmark datasets, considering six performance evaluation metrics. Comparative analyses with seven state-of-the-art methods are conducted. Furthermore, a stability analysis of DLRLA is performed across all datasets and evaluation metrics, showcasing its robustness and consistency.
期刊介绍:
Artificial Intelligence (AI) is pivotal in driving the fourth industrial revolution, witnessing remarkable advancements across various machine learning methodologies. AI techniques have become indispensable tools for practicing engineers, enabling them to tackle previously insurmountable challenges. Engineering Applications of Artificial Intelligence serves as a global platform for the swift dissemination of research elucidating the practical application of AI methods across all engineering disciplines. Submitted papers are expected to present novel aspects of AI utilized in real-world engineering applications, validated using publicly available datasets to ensure the replicability of research outcomes. Join us in exploring the transformative potential of AI in engineering.