{"title":"Multi-factor matching method for basic information of science and technology experts based on Web mining","authors":"Pei Zhou, Quanyin Zhu","doi":"10.1109/ICSESS.2012.6269567","DOIUrl":null,"url":null,"abstract":"The accuracy rate of information extracting by Web mining is not high because of the diversity and complexity of Web page. In order to increase the accuracy rate of information extracting by Web mining for building the science and technology basic information system, a novel multi-factor matching is proposed in this paper. The proposed method integrates the position of every word among the keywords corpus in normalized text and the multi-factor matching method between keywords corpus and normalized text which extracted from Web page by URL. The extracted results include the name, sex, birth, hometown and professional title of science and technology experts respectively. Experiments show that the accuracy rates obtain 95.64 percent and the recall rates achieve 99.69 percent respectively. The results show as by proposed method can satisfied the application requirements.","PeriodicalId":205738,"journal":{"name":"2012 IEEE International Conference on Computer Science and Automation Engineering","volume":"78 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2012-06-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2012 IEEE International Conference on Computer Science and Automation Engineering","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICSESS.2012.6269567","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
The accuracy rate of information extracting by Web mining is not high because of the diversity and complexity of Web page. In order to increase the accuracy rate of information extracting by Web mining for building the science and technology basic information system, a novel multi-factor matching is proposed in this paper. The proposed method integrates the position of every word among the keywords corpus in normalized text and the multi-factor matching method between keywords corpus and normalized text which extracted from Web page by URL. The extracted results include the name, sex, birth, hometown and professional title of science and technology experts respectively. Experiments show that the accuracy rates obtain 95.64 percent and the recall rates achieve 99.69 percent respectively. The results show as by proposed method can satisfied the application requirements.