{"title":"噪声阿拉伯文献的作者识别","authors":"S. Bourib, H. Sayoud","doi":"10.1109/CoDIT.2018.8394885","DOIUrl":null,"url":null,"abstract":"In the present research work, we deal with the problem of authorship attribution of Ancient Arabic Philosophers. For that purpose, we conducted several authorship attribution experiments applied to different noise Arabic text. A special dataset, called “A4P” (Authorship Attribution for Ancient Arabic Philosophers), has been constructed by extracting texts from the books of 5 Ancient Arabic Philosophers, where the genre and the topic are similar. In our approach two types of features were employed; character N-grams and words and several classifiers are used, namely: Support Vector Machines, Multi Layer Perceptron, Linear Regression, Stamatatos distance and Manhattan distance. The obtained results show that the failure limit and classification performances depend on the used features, the classification technique and the level of noise. In the overall the performances of the proposed techniques are quite interesting by showing the effect of noise on authorship attribution.","PeriodicalId":128011,"journal":{"name":"2018 5th International Conference on Control, Decision and Information Technologies (CoDIT)","volume":"18 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2018-04-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"Author Identification on Noise Arabic Documents\",\"authors\":\"S. Bourib, H. Sayoud\",\"doi\":\"10.1109/CoDIT.2018.8394885\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In the present research work, we deal with the problem of authorship attribution of Ancient Arabic Philosophers. For that purpose, we conducted several authorship attribution experiments applied to different noise Arabic text. A special dataset, called “A4P” (Authorship Attribution for Ancient Arabic Philosophers), has been constructed by extracting texts from the books of 5 Ancient Arabic Philosophers, where the genre and the topic are similar. In our approach two types of features were employed; character N-grams and words and several classifiers are used, namely: Support Vector Machines, Multi Layer Perceptron, Linear Regression, Stamatatos distance and Manhattan distance. The obtained results show that the failure limit and classification performances depend on the used features, the classification technique and the level of noise. In the overall the performances of the proposed techniques are quite interesting by showing the effect of noise on authorship attribution.\",\"PeriodicalId\":128011,\"journal\":{\"name\":\"2018 5th International Conference on Control, Decision and Information Technologies (CoDIT)\",\"volume\":\"18 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2018-04-10\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2018 5th International Conference on Control, Decision and Information Technologies (CoDIT)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/CoDIT.2018.8394885\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2018 5th International Conference on Control, Decision and Information Technologies (CoDIT)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/CoDIT.2018.8394885","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
In the present research work, we deal with the problem of authorship attribution of Ancient Arabic Philosophers. For that purpose, we conducted several authorship attribution experiments applied to different noise Arabic text. A special dataset, called “A4P” (Authorship Attribution for Ancient Arabic Philosophers), has been constructed by extracting texts from the books of 5 Ancient Arabic Philosophers, where the genre and the topic are similar. In our approach two types of features were employed; character N-grams and words and several classifiers are used, namely: Support Vector Machines, Multi Layer Perceptron, Linear Regression, Stamatatos distance and Manhattan distance. The obtained results show that the failure limit and classification performances depend on the used features, the classification technique and the level of noise. In the overall the performances of the proposed techniques are quite interesting by showing the effect of noise on authorship attribution.