HTML Pattern Generator--Automatic Data Extraction from Web Pages

2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing Pub Date : 2006-09-26 DOI:10.1109/SYNASC.2006.43

M. Cosulschi, A. Giurca, Bogdan Udrescu, N. Constantinescu, M. Gabroveanu

引用次数: 3

Abstract

Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved

查看原文本刊更多论文

HTML模式生成器——从网页中自动提取数据

现有的从HTML文档中提取信息的方法包括人工方法、监督学习和自动技术。手工方法具有较高的查全率和查全率，但难以适用于大量的页数。监督式学习涉及人类互动，以创造积极和消极的样本。自动化技术受益于较少的人力，但它们在检索信息方面不是高度可靠的

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing

自引率

0.00%

发文量