{"title":"基于类别路径信息的大规模层次文本分类窄带方法的改进","authors":"Heung-Seon Oh, Yuchul Jung","doi":"10.1633/JISTaP.2017.5.3.3","DOIUrl":null,"url":null,"abstract":"The narrow-down approach, separately composed of search and classification stages, is an effective way of dealing with large-scale hierarchical text classification. Recent approaches introduce methods of incorporating global, local, and path information extracted from web taxonomies in the classification stage. Meanwhile, in the case of utilizing path information, there have been few efforts to address existing limitations and develop more sophisticated methods. In this paper, we propose an expansion method to effectively exploit category path information based on the observation that the existing method is exposed to a term mismatch problem and low discrimination power due to insufficient path information. The key idea of our method is to utilize relevant information not presented on category paths by adding more useful words. We evaluate the effectiveness of our method on state-of-the art narrow-down methods and report the results with in-depth analysis.","PeriodicalId":37582,"journal":{"name":"Journal of Information Science Theory and Practice","volume":"17 1","pages":"31-47"},"PeriodicalIF":0.0000,"publicationDate":"2017-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Enhancing the Narrow-down Approach to Large-scale Hierarchical Text Classification with Category Path Information\",\"authors\":\"Heung-Seon Oh, Yuchul Jung\",\"doi\":\"10.1633/JISTaP.2017.5.3.3\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The narrow-down approach, separately composed of search and classification stages, is an effective way of dealing with large-scale hierarchical text classification. Recent approaches introduce methods of incorporating global, local, and path information extracted from web taxonomies in the classification stage. Meanwhile, in the case of utilizing path information, there have been few efforts to address existing limitations and develop more sophisticated methods. In this paper, we propose an expansion method to effectively exploit category path information based on the observation that the existing method is exposed to a term mismatch problem and low discrimination power due to insufficient path information. The key idea of our method is to utilize relevant information not presented on category paths by adding more useful words. We evaluate the effectiveness of our method on state-of-the art narrow-down methods and report the results with in-depth analysis.\",\"PeriodicalId\":37582,\"journal\":{\"name\":\"Journal of Information Science Theory and Practice\",\"volume\":\"17 1\",\"pages\":\"31-47\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2017-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Journal of Information Science Theory and Practice\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1633/JISTaP.2017.5.3.3\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q3\",\"JCRName\":\"Social Sciences\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Information Science Theory and Practice","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1633/JISTaP.2017.5.3.3","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"Social Sciences","Score":null,"Total":0}
Enhancing the Narrow-down Approach to Large-scale Hierarchical Text Classification with Category Path Information
The narrow-down approach, separately composed of search and classification stages, is an effective way of dealing with large-scale hierarchical text classification. Recent approaches introduce methods of incorporating global, local, and path information extracted from web taxonomies in the classification stage. Meanwhile, in the case of utilizing path information, there have been few efforts to address existing limitations and develop more sophisticated methods. In this paper, we propose an expansion method to effectively exploit category path information based on the observation that the existing method is exposed to a term mismatch problem and low discrimination power due to insufficient path information. The key idea of our method is to utilize relevant information not presented on category paths by adding more useful words. We evaluate the effectiveness of our method on state-of-the art narrow-down methods and report the results with in-depth analysis.
期刊介绍:
The Journal of Information Science Theory and Practice (JISTaP) is an international journal that aims at publishing original studies, review papers and brief communications on information science theory and practice. The journal provides an international forum for practical as well as theoretical research in the interdisciplinary areas of information science, such as information processing and management, knowledge organization, scholarly communication and bibliometrics. To foster scholarly communication among researchers and practitioners of library and information science around the globe, JISTaP offers a no-fee open access publishing venue where a team of dedicated editors, reviewers and staff members volunteer their services to ensure rapid dissemination and communication of scholarly works that make significant contributions. In a modern society, where information production and consumption grow at an astronomical rate, the science of information management, organization, and analysis is invaluable in effective utilization of information. The key objective of the journal is to foster research that can contribute to advancements and innovations in the theory and practice of information and library science so as to promote timely application of the findings from scientific investigations to everyday life. Recognizing the importance of the global perspective with understanding of region-specific issues, JISTaP encourages submissions of manuscripts that discuss global implications of regional findings as well as regional implications of global findings.