Compression-based arabic text classification

2014 IEEE/ACS 11th International Conference on Computer Systems and Applications (AICCSA) Pub Date : 2014-11-01 DOI:10.1109/AICCSA.2014.7073253

Haneen Ta'amneh, Ehsan Abu Keshek, M. B. Issa, M. Al-Ayyoub, Y. Jararweh

引用次数: 16

Abstract

Text classification (TC) is one of the fundamental problems in text mining. Plenty of works exist on TC with interesting approaches and excellent results; however, most of these works follow a word-based approach for feature extraction. In this work, we are interested in an alternative (byte-based or character-based) approach known as compression-based TC (CTC). CTC has been used for some languages such as English and Portuguese and it is shown to have certain advantages/ disadvantages compared with word-based approaches. This work applies CTC on the Arabic language with the purpose of investigating whether these advantages/disadvantages exists for the Arabic language as well. The results are encouraging as they show the viability of using CTC for Arabic TC.

查看原文本刊更多论文

基于压缩的阿拉伯语文本分类

文本分类(TC)是文本挖掘的基本问题之一。关于TC的研究有很多，方法有趣，结果也很好;然而，这些工作大多采用基于词的方法进行特征提取。在这项工作中，我们对另一种(基于字节或基于字符的)方法感兴趣，这种方法称为基于压缩的TC (CTC)。CTC已被用于一些语言，如英语和葡萄牙语，与基于单词的方法相比，它显示出一定的优点/缺点。这项工作将CTC应用于阿拉伯语，目的是调查阿拉伯语是否也存在这些优势/劣势。结果令人鼓舞，因为它们显示了将CTC用于阿拉伯语TC的可行性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2014 IEEE/ACS 11th International Conference on Computer Systems and Applications (AICCSA)

自引率

0.00%

发文量