Web data extraction techniques: A review

2016 World Conference on Futuristic Trends in Research and Innovation for Social Welfare (Startup Conclave) Pub Date : 2016-02-01 DOI:10.1109/STARTUP.2016.7583910

N. V. Kamanwar, S. Kale

引用次数: 8

Abstract

Web data extraction is the process of extracting user required information from websites. The web document contains data which is not in structured format. From the word web data extraction, we mean the extraction of data that is present in the web documents in HTML format. Then removing the unwanted stuff such as tags, advertisements, videos and so on. Then learning the information or patterns or features present in that data. Today, most researchers uses web data extractors because the internet contains huge data which makes the process of manual information extraction from the web documents complicated. In this paper, we have studied about different techniques for data extraction used by different authors that takes the user required data from a set of web pages. A comparative analysis of web data extraction techniques is given.

查看原文本刊更多论文

Web数据提取技术综述

Web数据提取是从网站中提取用户所需信息的过程。web文档包含非结构化格式的数据。从web数据提取这个词中，我们的意思是提取以HTML格式存在于web文档中的数据。然后删除不需要的东西，如标签，广告，视频等。然后学习数据中存在的信息、模式或特征。目前，大多数研究人员使用网络数据提取器，因为网络中包含大量的数据，这使得从网络文档中手动提取信息的过程变得复杂。在本文中，我们研究了不同作者使用的不同数据提取技术，这些技术从一组网页中获取用户所需的数据。对网络数据提取技术进行了比较分析。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2016 World Conference on Futuristic Trends in Research and Innovation for Social Welfare (Startup Conclave)

自引率

0.00%

发文量