Identifying bioentity recognition errors of rule-based text-mining systems

2008 Third International Conference on Digital Information Management Pub Date : 2008-11-01 DOI:10.1109/ICDIM.2008.4746791

Francisco M. Couto, Tiago Grego, Hugo P. Bastos, Catia Pesquita, Rafael P. Torres Jiménez, Pablo Sánchez, Leandro Pascual, C. Blaschke

{"title":"Identifying bioentity recognition errors of rule-based text-mining systems","authors":"Francisco M. Couto, Tiago Grego, Hugo P. Bastos, Catia Pesquita, Rafael P. Torres Jiménez, Pablo Sánchez, Leandro Pascual, C. Blaschke","doi":"10.1109/ICDIM.2008.4746791","DOIUrl":null,"url":null,"abstract":"An important research topic in Bioinformatics involves the exploration of vast amounts of biological and biomedical scientific literature (BioLiterature). Over the last few decades, text-mining systems have exploited this BioLiterature to reduce the time spent by researchers in its analysis. However, state-of-the-art approaches are still far from reaching performance levels acceptable by curators, and below the performance obtained in other domains, such as personal name recognition or news text. To achieve high levels of performance, it is essential that text mining tools effectively recognize bioentities present in BioLiterature. This paper presents FIBRE (Filtering Bioentity Recognition Errors), a system for automatically filtering mis annotations generated by rule-based systems that automatically recognize bioentities in BioLiterature. FIBRE aims at using different sets of automatically generated annotations to identify the main features that characterize an annotation of being of a certain type. These features are then used to filter mis annotations using a confidence threshold. The assessment of FIBRE was performed on a set of more than 17,000 documents, previously annotated by Text Detective, a state-of-the-art rule-based name bioentity recognition system. Curators evaluated the gene annotations given by Text Detective that FIBRE classified as non-gene annotations, and we found that FIBRE was able to filter with a precision above 92% more than 600 mis annotations, requiring minimal human effort, which demonstrates the effectiveness of FIBRE in a realistic scenario.","PeriodicalId":415013,"journal":{"name":"2008 Third International Conference on Digital Information Management","volume":"46 3-4","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2008-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"4","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2008 Third International Conference on Digital Information Management","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDIM.2008.4746791","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 4

Abstract

An important research topic in Bioinformatics involves the exploration of vast amounts of biological and biomedical scientific literature (BioLiterature). Over the last few decades, text-mining systems have exploited this BioLiterature to reduce the time spent by researchers in its analysis. However, state-of-the-art approaches are still far from reaching performance levels acceptable by curators, and below the performance obtained in other domains, such as personal name recognition or news text. To achieve high levels of performance, it is essential that text mining tools effectively recognize bioentities present in BioLiterature. This paper presents FIBRE (Filtering Bioentity Recognition Errors), a system for automatically filtering mis annotations generated by rule-based systems that automatically recognize bioentities in BioLiterature. FIBRE aims at using different sets of automatically generated annotations to identify the main features that characterize an annotation of being of a certain type. These features are then used to filter mis annotations using a confidence threshold. The assessment of FIBRE was performed on a set of more than 17,000 documents, previously annotated by Text Detective, a state-of-the-art rule-based name bioentity recognition system. Curators evaluated the gene annotations given by Text Detective that FIBRE classified as non-gene annotations, and we found that FIBRE was able to filter with a precision above 92% more than 600 mis annotations, requiring minimal human effort, which demonstrates the effectiveness of FIBRE in a realistic scenario.

查看原文本刊更多论文

基于规则的文本挖掘系统生物实体识别错误的识别

生物信息学的一个重要研究课题是对大量生物和生物医学科学文献(BioLiterature)的探索。在过去的几十年里，文本挖掘系统利用这种生物文献来减少研究人员在分析中花费的时间。然而，最先进的方法仍远未达到策展人可接受的性能水平，并且低于其他领域(如个人姓名识别或新闻文本)所获得的性能。为了达到高水平的性能，文本挖掘工具有效识别生物文献中存在的生物实体是至关重要的。本文介绍了纤维(过滤生物实体识别错误)，一个自动过滤由基于规则的系统生成的错误注释的系统，该系统自动识别生物文献中的生物实体。FIBRE的目的是使用不同的自动生成注释集来识别某种类型注释的主要特征。然后使用这些特性使用置信度阈值过滤错误的注释。fiber的评估是在一组超过17,000份文件上进行的，这些文件之前由Text Detective(一种最先进的基于规则的名称生物实体识别系统)注释。管理员评估了文本侦探给出的基因注释，纤维将其归类为非基因注释，我们发现纤维能够以超过92%的精度过滤600多个错误注释，需要最少的人力，这证明了纤维在现实场景中的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2008 Third International Conference on Digital Information Management

自引率

0.00%

发文量