Automatic Labeling of Topics

2009 Ninth International Conference on Intelligent Systems Design and Applications Pub Date : 2009-11-30 DOI:10.1109/ISDA.2009.165

Davide Magatti, S. Calegari, D. Ciucci, Fabio Stella

引用次数: 69

Abstract

An algorithm for the automatic labeling of topics accordingly to a hierarchy is presented. Its main ingredients are a set of similarity measures and a set of topic labeling rules. The labeling rules are specifically designed to find the most agreed labels between the given topic and the hierarchy. The hierarchy is obtained from the Google Directory service, extracted via an ad-hoc developed software procedure and expanded through the use of the OpenOffice English Thesaurus. The performance of the proposed algorithm is investigated by using a document corpus consisting of 33,801 documents and a dictionary consisting of 111,795 words. The results are encouraging, while particularly interesting and significant labeling cases emerged

查看原文本刊更多论文

自动标注主题

提出了一种基于层次结构的主题自动标注算法。它的主要成分是一组相似度度量和一组主题标注规则。标记规则专门用于找到给定主题和层次结构之间最一致的标签。层次结构从谷歌Directory服务获得，通过特别开发的软件过程提取，并通过使用OpenOffice English Thesaurus进行扩展。通过使用包含33,801个文档的文档语料库和包含111,795个单词的字典来研究该算法的性能。结果令人鼓舞，同时出现了特别有趣和重要的标签案例

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2009 Ninth International Conference on Intelligent Systems Design and Applications

自引率

0.00%

发文量