{"title":"The impact of arabic inter-character proximity and similarity on spell-checking","authors":"Hicham Gueddah, Abdallah Yousfi","doi":"10.1109/SITA.2013.6560811","DOIUrl":null,"url":null,"abstract":"Following a statistical study carried out on the typographical errors committed when typing documents in Arabic language, it was found that most of these typos are character permutation errors, accounting for 65% of overall errors. Similarly, in the aim to overcome this situation and while analyzing such errors, it turned out that character permutation resulting in a misspelled word can be imputed either to character proximity on an Arabic keyboard or to calligraphic similarity between such Arabic characters. In order to remedy this problem, we suggest, in this article, that a measurement of proximity and similarity between Arabic characters be integrated into Levenshtein algorithm, in the aim of enhancing the suggestions and the scheduling of the solutions returned by the spelling correction. The experimental outcomes are very satisfactory and attest of the necessity of integrating inter-character proximity and similarity measures for Arabic within spell-checking systems.","PeriodicalId":145244,"journal":{"name":"2013 8th International Conference on Intelligent Systems: Theories and Applications (SITA)","volume":"104 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2013-05-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"9","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2013 8th International Conference on Intelligent Systems: Theories and Applications (SITA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SITA.2013.6560811","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 9
Abstract
Following a statistical study carried out on the typographical errors committed when typing documents in Arabic language, it was found that most of these typos are character permutation errors, accounting for 65% of overall errors. Similarly, in the aim to overcome this situation and while analyzing such errors, it turned out that character permutation resulting in a misspelled word can be imputed either to character proximity on an Arabic keyboard or to calligraphic similarity between such Arabic characters. In order to remedy this problem, we suggest, in this article, that a measurement of proximity and similarity between Arabic characters be integrated into Levenshtein algorithm, in the aim of enhancing the suggestions and the scheduling of the solutions returned by the spelling correction. The experimental outcomes are very satisfactory and attest of the necessity of integrating inter-character proximity and similarity measures for Arabic within spell-checking systems.