Automatic classification of documents by formality

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010) Pub Date : 2010-09-30 DOI:10.1109/NLPKE.2010.5587767

Fadi Abu Sheikha, D. Inkpen

引用次数: 30

Abstract

This paper addresses the task of classifying documents into formal or informal style. We studied the main characteristics of each style in order to choose features that allowed us to train classifiers that can distinguish between the two styles. We built our data set by collecting documents for both styles, from different sources. We tested several classification algorithms, namely Decision Trees, Naïve Bayes, and Support Vector Machines, to choose the classifier that leads to the best classification results. We performed attribute selection in order to determine the contribution of each feature to our model.

查看原文本刊更多论文

按形式自动分类文件

本文解决了将文档分类为正式或非正式风格的任务。我们研究了每种风格的主要特征，以便选择能够让我们训练能够区分两种风格的分类器的特征。我们通过收集来自不同来源的两种风格的文档来构建数据集。我们测试了几种分类算法，即决策树，Naïve贝叶斯和支持向量机，以选择导致最佳分类结果的分类器。为了确定每个特征对模型的贡献，我们执行了属性选择。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering(NLPKE-2010)

自引率

0.00%

发文量