Multilingual Resources for Entity Extraction

NER@ACL Pub Date : 2003-07-12 DOI:10.3115/1119384.1119391

S. Strassel, A. Mitchell

引用次数: 18

Abstract

Progress in human language technology requires increasing amounts of data and annotation in a growing variety of languages. Research in Named Entity extraction is no exception. Linguistic Data Consortium is creating annotated corpora to support information extraction in English, Chinese, Arabic, and other languages for a variety of US Government-sponsored programs. This paper covers the scope of annotation and research tasks within these programs, describes some of the challenges of multilingual corpus development for entity extraction, and concludes with a description of the corpora developed to support this research.

查看原文本刊更多论文

实体抽取的多语言资源

人类语言技术的进步需要越来越多的数据和各种语言的注释。命名实体提取的研究也不例外。语言数据联盟正在创建带注释的语料库，以支持英语、中文、阿拉伯语和其他语言的信息提取，用于各种美国政府资助的项目。本文涵盖了这些项目中的注释范围和研究任务，描述了用于实体提取的多语言语料库开发的一些挑战，并最后描述了为支持本研究而开发的语料库。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

NER@ACL

自引率

0.00%

发文量