Combining information extraction and human computing for crowdsourced knowledge acquisition

2014 IEEE 30th International Conference on Data Engineering Pub Date : 2014-03-01 DOI:10.1109/ICDE.2014.6816717

Sarath Kumar Kondreddi, P. Triantafillou, G. Weikum

{"title":"Combining information extraction and human computing for crowdsourced knowledge acquisition","authors":"Sarath Kumar Kondreddi, P. Triantafillou, G. Weikum","doi":"10.1109/ICDE.2014.6816717","DOIUrl":null,"url":null,"abstract":"Automatic information extraction (IE) enables the construction of very large knowledge bases (KBs), with relational facts on millions of entities from text corpora and Web sources. However, such KBs contain errors and they are far from being complete. This motivates the need for exploiting human intelligence and knowledge using crowd-based human computing (HC) for assessing the validity of facts and for gathering additional knowledge. This paper presents a novel system architecture, called Higgins, which shows how to effectively integrate an IE engine and a HC engine. Higgins generates game questions where players choose or fill in missing relations for subject-relation-object triples. For generating multiple-choice answer candidates, we have constructed a large dictionary of entity names and relational phrases, and have developed specifically designed statistical language models for phrase relatedness. To this end, we combine semantic resources like WordNet, ConceptNet, and others with statistics derived from a large Web corpus. We demonstrate the effectiveness of Higgins for knowledge acquisition by crowdsourced gathering of relationships between characters in narrative descriptions of movies and books.","PeriodicalId":159130,"journal":{"name":"2014 IEEE 30th International Conference on Data Engineering","volume":"36 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"53","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2014 IEEE 30th International Conference on Data Engineering","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDE.2014.6816717","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 53

Abstract

Automatic information extraction (IE) enables the construction of very large knowledge bases (KBs), with relational facts on millions of entities from text corpora and Web sources. However, such KBs contain errors and they are far from being complete. This motivates the need for exploiting human intelligence and knowledge using crowd-based human computing (HC) for assessing the validity of facts and for gathering additional knowledge. This paper presents a novel system architecture, called Higgins, which shows how to effectively integrate an IE engine and a HC engine. Higgins generates game questions where players choose or fill in missing relations for subject-relation-object triples. For generating multiple-choice answer candidates, we have constructed a large dictionary of entity names and relational phrases, and have developed specifically designed statistical language models for phrase relatedness. To this end, we combine semantic resources like WordNet, ConceptNet, and others with statistics derived from a large Web corpus. We demonstrate the effectiveness of Higgins for knowledge acquisition by crowdsourced gathering of relationships between characters in narrative descriptions of movies and books.

查看原文本刊更多论文

信息提取与人工计算相结合的众包知识获取

自动信息提取(IE)支持构建非常大的知识库(KBs)，其中包含来自文本语料库和Web源的数百万个实体的关系事实。但是，这样的KBs有错误，而且还远远不够完整。这激发了利用基于人群的人类计算(HC)来评估事实的有效性和收集额外知识的人类智能和知识的需求。本文提出了一种新的系统架构，称为Higgins，它展示了如何有效地集成IE引擎和HC引擎。希金斯生成了一些游戏问题，让玩家选择或填写主体-关系-客体三元组中缺失的关系。为了生成选择题候选答案，我们构建了一个大型实体名称和关系短语字典，并开发了专门设计的短语相关性统计语言模型。为此，我们将WordNet、ConceptNet等语义资源与来自大型Web语料库的统计数据结合起来。我们通过对电影和书籍叙事描述中人物关系的众包收集，证明了希金斯理论在知识获取方面的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2014 IEEE 30th International Conference on Data Engineering

自引率

0.00%

发文量