{"title":"基于大数据和机器学习的对话文本特征提取","authors":"Xueli Liu, Hua Zhang, Yue Cheng","doi":"10.4018/ijwltt.337602","DOIUrl":null,"url":null,"abstract":"In this article, a dialogue text feature extraction model based on big data and machine learning is constructed, which transforms the high-dimensional space of text features into the low-dimensional space that is easy to process, so that the best feature words can be selected to represent the document set. Tests show that in most cases, the classification accuracy of this model is higher than 88%, and the recall rate is higher than 85%, thus achieving the goal of higher classification accuracy with less computation. When extracting the features of dialogue texts, there is no need for preprocessing, just count the data such as lexical composition, sentence length and sentence-to-sentence relationship of the target text, and make linear analysis to obtain key indicators and weights. Based on this, the classification model can achieve good results, thus effectively reducing the workload and computation of text classification.","PeriodicalId":506438,"journal":{"name":"International Journal of Web-Based Learning and Teaching Technologies","volume":"63 3","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2024-02-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Feature Extraction of Dialogue Text Based on Big Data and Machine Learning\",\"authors\":\"Xueli Liu, Hua Zhang, Yue Cheng\",\"doi\":\"10.4018/ijwltt.337602\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this article, a dialogue text feature extraction model based on big data and machine learning is constructed, which transforms the high-dimensional space of text features into the low-dimensional space that is easy to process, so that the best feature words can be selected to represent the document set. Tests show that in most cases, the classification accuracy of this model is higher than 88%, and the recall rate is higher than 85%, thus achieving the goal of higher classification accuracy with less computation. When extracting the features of dialogue texts, there is no need for preprocessing, just count the data such as lexical composition, sentence length and sentence-to-sentence relationship of the target text, and make linear analysis to obtain key indicators and weights. Based on this, the classification model can achieve good results, thus effectively reducing the workload and computation of text classification.\",\"PeriodicalId\":506438,\"journal\":{\"name\":\"International Journal of Web-Based Learning and Teaching Technologies\",\"volume\":\"63 3\",\"pages\":\"\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2024-02-07\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"International Journal of Web-Based Learning and Teaching Technologies\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.4018/ijwltt.337602\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"International Journal of Web-Based Learning and Teaching Technologies","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.4018/ijwltt.337602","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Feature Extraction of Dialogue Text Based on Big Data and Machine Learning
In this article, a dialogue text feature extraction model based on big data and machine learning is constructed, which transforms the high-dimensional space of text features into the low-dimensional space that is easy to process, so that the best feature words can be selected to represent the document set. Tests show that in most cases, the classification accuracy of this model is higher than 88%, and the recall rate is higher than 85%, thus achieving the goal of higher classification accuracy with less computation. When extracting the features of dialogue texts, there is no need for preprocessing, just count the data such as lexical composition, sentence length and sentence-to-sentence relationship of the target text, and make linear analysis to obtain key indicators and weights. Based on this, the classification model can achieve good results, thus effectively reducing the workload and computation of text classification.