Jose L. Hurtado, Napat Taweewitchakreeya, Xingquan Zhu
{"title":"这篇论文是谁写的?使用文体学特征学习作者身份去识别","authors":"Jose L. Hurtado, Napat Taweewitchakreeya, Xingquan Zhu","doi":"10.1109/IRI.2014.7051981","DOIUrl":null,"url":null,"abstract":"In this paper, we propose to combine stylometric features and neural networks for authorship de-identification. Our research mainly focuses on scientific publications, because scholarly journals are publicly available with plenty of labeled data to learn an author's style or traits. The main challenge of authorship de-identification is to identify features which can properly capture an author's writing style. In the proposed design, we choose a combination of stylometric features, including lexical, syntactic, structural and content-specific features, to represent each author's style and use them to build classification models. We manually collect publications from computer science and biomedicine domains and validate our designs by using a number of classification methods. Our experiments show that among four well-known classifiers, Multilayer Perceptron (MLP) classifiers achieve the best performance for authorship de-identification.","PeriodicalId":360013,"journal":{"name":"Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration (IEEE IRI 2014)","volume":"7 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"8","resultStr":"{\"title\":\"Who wrote this paper? Learning for authorship de-identification using stylometric featuress\",\"authors\":\"Jose L. Hurtado, Napat Taweewitchakreeya, Xingquan Zhu\",\"doi\":\"10.1109/IRI.2014.7051981\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper, we propose to combine stylometric features and neural networks for authorship de-identification. Our research mainly focuses on scientific publications, because scholarly journals are publicly available with plenty of labeled data to learn an author's style or traits. The main challenge of authorship de-identification is to identify features which can properly capture an author's writing style. In the proposed design, we choose a combination of stylometric features, including lexical, syntactic, structural and content-specific features, to represent each author's style and use them to build classification models. We manually collect publications from computer science and biomedicine domains and validate our designs by using a number of classification methods. Our experiments show that among four well-known classifiers, Multilayer Perceptron (MLP) classifiers achieve the best performance for authorship de-identification.\",\"PeriodicalId\":360013,\"journal\":{\"name\":\"Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration (IEEE IRI 2014)\",\"volume\":\"7 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2014-08-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"8\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration (IEEE IRI 2014)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/IRI.2014.7051981\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2014 IEEE 15th International Conference on Information Reuse and Integration (IEEE IRI 2014)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IRI.2014.7051981","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Who wrote this paper? Learning for authorship de-identification using stylometric featuress
In this paper, we propose to combine stylometric features and neural networks for authorship de-identification. Our research mainly focuses on scientific publications, because scholarly journals are publicly available with plenty of labeled data to learn an author's style or traits. The main challenge of authorship de-identification is to identify features which can properly capture an author's writing style. In the proposed design, we choose a combination of stylometric features, including lexical, syntactic, structural and content-specific features, to represent each author's style and use them to build classification models. We manually collect publications from computer science and biomedicine domains and validate our designs by using a number of classification methods. Our experiments show that among four well-known classifiers, Multilayer Perceptron (MLP) classifiers achieve the best performance for authorship de-identification.