HTML模式生成器——从网页中自动提取数据

2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing Pub Date : 2006-09-26 DOI:10.1109/SYNASC.2006.43

M. Cosulschi, A. Giurca, Bogdan Udrescu, N. Constantinescu, M. Gabroveanu

{"title":"HTML模式生成器——从网页中自动提取数据","authors":"M. Cosulschi, A. Giurca, Bogdan Udrescu, N. Constantinescu, M. Gabroveanu","doi":"10.1109/SYNASC.2006.43","DOIUrl":null,"url":null,"abstract":"Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved","PeriodicalId":309740,"journal":{"name":"2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing","volume":"12 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2006-09-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":"{\"title\":\"HTML Pattern Generator--Automatic Data Extraction from Web Pages\",\"authors\":\"M. Cosulschi, A. Giurca, Bogdan Udrescu, N. Constantinescu, M. Gabroveanu\",\"doi\":\"10.1109/SYNASC.2006.43\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved\",\"PeriodicalId\":309740,\"journal\":{\"name\":\"2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing\",\"volume\":\"12 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2006-09-26\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"3\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SYNASC.2006.43\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SYNASC.2006.43","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 3

摘要

现有的从HTML文档中提取信息的方法包括人工方法、监督学习和自动技术。手工方法具有较高的查全率和查全率，但难以适用于大量的页数。监督式学习涉及人类互动，以创造积极和消极的样本。自动化技术受益于较少的人力，但它们在检索信息方面不是高度可靠的

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

HTML Pattern Generator--Automatic Data Extraction from Web Pages

Existing methods of information extraction from HTML documents include manual approach, supervised learning and automatic techniques. The manual method has high precision and recall values but it is difficult to apply it for large number of pages. Supervised learning involves human interaction to create positive and negative samples. Automatic techniques benefit from less human effort but they are not highly reliable regarding the information retrieved

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2006 Eighth International Symposium on Symbolic and Numeric Algorithms for Scientific Computing

自引率

0.00%

发文量