{"title":"基于LDA主题模型的短文本分类","authors":"Qiuxing Chen, Lixiu Yao, Jie Yang","doi":"10.1109/ICALIP.2016.7846525","DOIUrl":null,"url":null,"abstract":"As the rapid development of computer technology and network communication, short text data has increased enormously. Classifying the short text snippets is a great challenge to due to its less semantic information and high sparseness. In this paper, we proposed an improved short text classification method based on Latent Dirichlet Allocation topic model and K-Nearest Neighbor algorithm. The generated probabilistic topics help both make the texts more semantic-focused and reduce the sparseness. In addition, we present a novel topic similarity measure method with the topic-word matrix and the relationship of the discriminative terms between two short texts. A short text dataset for experiment validation is constructed by crawling the posts from Sina News website. The extensive and comparable experimental results obtained show the effectiveness of our proposed method.","PeriodicalId":184170,"journal":{"name":"2016 International Conference on Audio, Language and Image Processing (ICALIP)","volume":"28 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2016-07-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"63","resultStr":"{\"title\":\"Short text classification based on LDA topic model\",\"authors\":\"Qiuxing Chen, Lixiu Yao, Jie Yang\",\"doi\":\"10.1109/ICALIP.2016.7846525\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"As the rapid development of computer technology and network communication, short text data has increased enormously. Classifying the short text snippets is a great challenge to due to its less semantic information and high sparseness. In this paper, we proposed an improved short text classification method based on Latent Dirichlet Allocation topic model and K-Nearest Neighbor algorithm. The generated probabilistic topics help both make the texts more semantic-focused and reduce the sparseness. In addition, we present a novel topic similarity measure method with the topic-word matrix and the relationship of the discriminative terms between two short texts. A short text dataset for experiment validation is constructed by crawling the posts from Sina News website. The extensive and comparable experimental results obtained show the effectiveness of our proposed method.\",\"PeriodicalId\":184170,\"journal\":{\"name\":\"2016 International Conference on Audio, Language and Image Processing (ICALIP)\",\"volume\":\"28 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2016-07-11\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"63\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2016 International Conference on Audio, Language and Image Processing (ICALIP)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICALIP.2016.7846525\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2016 International Conference on Audio, Language and Image Processing (ICALIP)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICALIP.2016.7846525","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Short text classification based on LDA topic model
As the rapid development of computer technology and network communication, short text data has increased enormously. Classifying the short text snippets is a great challenge to due to its less semantic information and high sparseness. In this paper, we proposed an improved short text classification method based on Latent Dirichlet Allocation topic model and K-Nearest Neighbor algorithm. The generated probabilistic topics help both make the texts more semantic-focused and reduce the sparseness. In addition, we present a novel topic similarity measure method with the topic-word matrix and the relationship of the discriminative terms between two short texts. A short text dataset for experiment validation is constructed by crawling the posts from Sina News website. The extensive and comparable experimental results obtained show the effectiveness of our proposed method.