Reference metadata extraction from scientific papers

2011 12th International Conference on Parallel and Distributed Computing, Applications and Technologies Pub Date : 2011-10-20 DOI:10.1109/PDCAT.2011.72

Zhixin Guo, Hai Jin

引用次数: 14

Abstract

Bibliographical information of scientific papers is of great value since the Science Citation Index is introduced to measure research impact. Most scientific documents available on the web are unstructured or semi-structured, and the automatic reference metadata extraction process becomes an important task. This paper describes a framework for automatic reference metadata extraction from scientific papers. Our system can extract title, author, journal, volume, year, and page from scientific papers in PDF. We utilize a document metadata knowledge base to guide the reference metadata extraction process. The experiment results show that our system achieves a high accuracy.

查看原文本刊更多论文

从科学论文中提取参考元数据

引入科学引文索引来衡量科研影响，使得科技论文的文献信息具有重要的价值。网络上的科学文献大多是非结构化或半结构化的，文献元数据的自动提取成为一项重要的任务。本文描述了一个自动提取科技论文参考元数据的框架。我们的系统可以从PDF格式的科学论文中提取标题、作者、期刊、卷数、年份和页面。我们利用文档元数据知识库来指导参考元数据的提取过程。实验结果表明，该系统达到了较高的精度。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2011 12th International Conference on Parallel and Distributed Computing, Applications and Technologies

自引率

0.00%

发文量