Identifying gene and protein names from biological texts

Computational Systems Bioinformatics. CSB2003. Proceedings of the 2003 IEEE Bioinformatics Conference. CSB2003 Pub Date : 2003-08-11 DOI:10.1109/CSB.2003.1227431

Weijian Xuan, S. Watson, H. Akil, F. Meng

引用次数: 7

Abstract

Extracting and identifying gene and protein names from literature is a critical step for mining functional information of genes and proteins. While extensive efforts have been devoted to this important task, most of them were aiming at extracting gene/protein name per se without paying much attention to associate the extracted name with existing gene and protein database entries. We developed a simple and efficient method to identify gene and protein names in literature using a combination of heuristic and statistical strategies. Our approach will map the extracted names to individual LocusLink entries thus enable the seamless integration of literature information with existing gene/protein databases. Evaluation on a test corpus shows that our method can achieve both high recall and precision. Our method exhibits good performance and can be used as a building block for large biomedical literature mining systems.

查看原文本刊更多论文

从生物学文本中识别基因和蛋白质名称

从文献中提取和识别基因和蛋白质的名称是挖掘基因和蛋白质功能信息的关键步骤。虽然在这一重要任务上已经投入了大量的努力，但大多数都是针对提取基因/蛋白质名称本身，而没有注意将提取的名称与现有的基因和蛋白质数据库条目相关联。我们开发了一种简单有效的方法来识别基因和蛋白质的名称在文献中使用启发式和统计策略的组合。我们的方法将提取的名称映射到单个LocusLink条目，从而实现文献信息与现有基因/蛋白质数据库的无缝集成。对测试语料库的评价表明，该方法具有较高的查全率和查准率。该方法具有良好的性能，可作为大型生物医学文献挖掘系统的基础。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Computational Systems Bioinformatics. CSB2003. Proceedings of the 2003 IEEE Bioinformatics Conference. CSB2003

自引率

0.00%

发文量