Text classification based on limited bibliographic metadata

2009 Fourth International Conference on Digital Information Management Pub Date : 2009-12-18 DOI:10.1109/ICDIM.2009.5356767

K. Denecke, T. Risse, Thomas Baehr

引用次数: 6

Abstract

In this paper, we introduce a method for categorizing digital items according to their topic, only relying on the document's metadata, such as author name and title information. The proposed approach is based on a set of lexical resources constructed for our purposes (e.g., journal titles, conference names) and on a traditional machine-learning classifier that assigns one category to each document based on identified core features. The system is evaluated on a real-world data set and the influence of different feature combinations and settings is studied. Although the available information is limited, the results show that the approach is capable to efficiently classify data items representing documents.

查看原文本刊更多论文

基于有限书目元数据的文本分类

在本文中，我们介绍了一种根据主题对数字项目进行分类的方法，该方法仅依赖于文档的元数据，如作者姓名和标题信息。提出的方法基于为我们的目的而构建的一组词汇资源(例如，期刊标题，会议名称)和传统的机器学习分类器，该分类器根据识别的核心特征为每个文档分配一个类别。在实际数据集上对系统进行了评估，并研究了不同特征组合和设置的影响。尽管可用的信息有限，但结果表明，该方法能够有效地对表示文档的数据项进行分类。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2009 Fourth International Conference on Digital Information Management

自引率

0.00%

发文量