Towards a high quality and web-scalable table search engine

KEYS '12 Pub Date : 2012-05-20 DOI:10.1145/2254736.2254738
Cong Yu
{"title":"Towards a high quality and web-scalable table search engine","authors":"Cong Yu","doi":"10.1145/2254736.2254738","DOIUrl":null,"url":null,"abstract":"For over a decade, a large number of studies have explored efficient mechanisms for finding relevant information from structured sources such as relational, semi-structured, and graph databases. While successful in its own right, keyword search over structured data has yet to gain wide spread adoption on the Web, and it is not because of the lack of structured data on the Web. In fact, the Web offers orders of magnitude more structured data than any offline data source: 14 billions tables can be gathered just by considering page content between the table tags. The main reason is that keyword search over structured data on the Web presents a unique set of challenges that are quite different from its non-Web counterparts, as well as different from searching over documents (where the search engines have excelled). In this talk, I will discuss those challenges and our approaches in addressing them at Google's WebTables project. In particular, I will present the table search engine, our initial effort toward building a high quality and scalable structured data search engine, along with our other efforts on managing structured data on the Web.","PeriodicalId":170987,"journal":{"name":"KEYS '12","volume":"43 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2012-05-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"7","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"KEYS '12","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/2254736.2254738","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 7

Abstract

For over a decade, a large number of studies have explored efficient mechanisms for finding relevant information from structured sources such as relational, semi-structured, and graph databases. While successful in its own right, keyword search over structured data has yet to gain wide spread adoption on the Web, and it is not because of the lack of structured data on the Web. In fact, the Web offers orders of magnitude more structured data than any offline data source: 14 billions tables can be gathered just by considering page content between the table tags. The main reason is that keyword search over structured data on the Web presents a unique set of challenges that are quite different from its non-Web counterparts, as well as different from searching over documents (where the search engines have excelled). In this talk, I will discuss those challenges and our approaches in addressing them at Google's WebTables project. In particular, I will present the table search engine, our initial effort toward building a high quality and scalable structured data search engine, along with our other efforts on managing structured data on the Web.
朝着高质量和网络可扩展的表搜索引擎
十多年来,大量的研究探索了从结构化来源(如关系数据库、半结构化数据库和图形数据库)中查找相关信息的有效机制。虽然结构化数据的关键字搜索本身就很成功,但它还没有在Web上得到广泛的采用,这并不是因为Web上缺乏结构化数据。事实上,Web提供的结构化数据比任何离线数据源都要多出数量级:仅考虑表标记之间的页面内容就可以收集140亿个表。主要原因是,在Web上对结构化数据进行关键字搜索呈现出一组独特的挑战,这些挑战与非Web对等项截然不同,也不同于对文档进行搜索(搜索引擎在这方面表现出色)。在这次演讲中,我将讨论这些挑战以及我们在Google的WebTables项目中解决这些挑战的方法。特别地,我将介绍表搜索引擎,这是我们为构建高质量和可伸缩的结构化数据搜索引擎所做的初步努力,以及我们在管理Web上的结构化数据方面所做的其他努力。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信