Review of web crawlers

Int. J. Knowl. Web Intell. Pub Date : 2014-10-01 DOI:10.1504/IJKWI.2014.065035

S. R. Sreeja, S. Chaudhari

引用次数: 5

Abstract

The web is a repository of large amount of data. Information available in the web is organised in the form of pages. Due to the presence of unlimited amount of information, searching and finding out appropriate information from the web is a task which needs expertise. Web crawlers are programmes that assist search engines by automating the task of visiting web pages and downloading their contents. They also help in ranking the downloaded web pages. Thus, the search engines can produce a list of web pages ordered by their relevance and can display this list as a result of the search. Crawling also helps to validate web pages, analyse them, notify about page-updation, visualise web pages and sometimes for collecting e-mail addresses for spam purposes. They can be of different types, each one using different strategies and techniques to crawl web pages. This paper presents a review of various types of web crawlers.

查看原文本刊更多论文

回顾网络爬虫

网络是大量数据的储存库。网络上可用的信息以页面的形式组织起来。由于信息的存在是无限的，从网络上搜索和找到合适的信息是一项需要专业知识的任务。网络爬虫是通过自动访问网页和下载其内容来协助搜索引擎的程序。它们还有助于对下载的网页进行排名。因此，搜索引擎可以产生一个网页列表，按其相关性排序，并可以显示这个列表作为搜索的结果。抓取还有助于验证网页、分析网页、通知网页更新、可视化网页，有时还用于收集电子邮件地址以用于垃圾邮件目的。它们可以是不同的类型，每个都使用不同的策略和技术来抓取网页。本文介绍了各种类型的网络爬虫。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Int. J. Knowl. Web Intell.

自引率

0.00%

发文量