MICHAEL:挖掘阿拉伯语方言识别的字符级模式(MADAR挑战)

WANLP@ACL 2019 Pub Date : 2019-08-01 DOI:10.18653/v1/W19-4627

Dhaou Ghoul, Gaël Lejeune

{"title":"MICHAEL:挖掘阿拉伯语方言识别的字符级模式(MADAR挑战)","authors":"Dhaou Ghoul, Gaël Lejeune","doi":"10.18653/v1/W19-4627","DOIUrl":null,"url":null,"abstract":"We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-processing free classification. More precisely, Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. This system achieved an official score (accuracy) of 53.25% with 1<=N<=3 but showed a much better result with character 4-grams (62.17% accuracy).","PeriodicalId":268163,"journal":{"name":"WANLP@ACL 2019","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)\",\"authors\":\"Dhaou Ghoul, Gaël Lejeune\",\"doi\":\"10.18653/v1/W19-4627\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-processing free classification. More precisely, Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. This system achieved an official score (accuracy) of 53.25% with 1<=N<=3 but showed a much better result with character 4-grams (62.17% accuracy).\",\"PeriodicalId\":268163,\"journal\":{\"name\":\"WANLP@ACL 2019\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-08-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"WANLP@ACL 2019\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.18653/v1/W19-4627\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"WANLP@ACL 2019","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/W19-4627","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

摘要

本文提出了一种基于MADAR旅行域方言识别(DID)的简易轻量级阿拉伯语方言自动识别方法MICHAEL。MICHAEL使用简单的字符级特征来执行预处理自由分类。更准确地说，从原始句子中提取的字符N-grams用于训练多项式朴素贝叶斯分类器。该系统在1<=N<=3时的官方得分(正确率)为53.25%，但在4克字符时的结果要好得多(正确率为62.17%)。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

MICHAEL: Mining Character-level Patterns for Arabic Dialect Identification (MADAR Challenge)

We present MICHAEL, a simple lightweight method for automatic Arabic Dialect Identification on the MADAR travel domain Dialect Identification (DID). MICHAEL uses simple character-level features in order to perform a pre-processing free classification. More precisely, Character N-grams extracted from the original sentences are used to train a Multinomial Naive Bayes classifier. This system achieved an official score (accuracy) of 53.25% with 1<=N<=3 but showed a much better result with character 4-grams (62.17% accuracy).

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

WANLP@ACL 2019

自引率

0.00%

发文量