A performance model for Forward XPath

2012 International Conference on High Performance Computing & Simulation (HPCS) Pub Date : 2012-07-02 DOI:10.1109/HPCSim.2012.6266979

M. Alrammal, G. Hains

引用次数: 0

Abstract

XML is a key standard for manipulating data on the Internet. However, querying large volume of XML data represents a bottleneck for several data intensive applications. Many modern applications require processing of massive streams of XML data, creating difficult technical challenges. Among these is the optimization of XPath query processing and accurate cost estimation for these queries when processed on a massive steam of XML data. In this paper, we present a novel performance prediction model which a priori estimates the cost of any Forward XPath structural in terms of space used and time spent. The model consists of (1) a lazy stream-querying algorithm LQ (2) a mathematical performance model (linear regression functions), and (3) a new selectivity estimation technique. Extensive experiments on both real and synthetic data sets show that our model achieves accuracy better than existing approaches. The resulting prototype supports the a priori design of efficient queries, as well as automatic query optimizations.

查看原文本刊更多论文

Forward XPath的性能模型

XML是在Internet上操作数据的关键标准。然而，查询大量XML数据对于一些数据密集型应用程序来说是一个瓶颈。许多现代应用程序需要处理大量XML数据流，这带来了困难的技术挑战。其中包括XPath查询处理的优化，以及在处理大量XML数据时对这些查询进行准确的成本估计。在本文中，我们提出了一种新的性能预测模型，该模型根据使用的空间和花费的时间先验地估计任何Forward XPath结构的成本。该模型包括:(1)延迟流查询算法LQ;(2)数学性能模型(线性回归函数);(3)一种新的选择性估计技术。在真实和合成数据集上的大量实验表明，我们的模型比现有的方法具有更好的精度。得到的原型支持高效查询的先验设计，以及自动查询优化。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2012 International Conference on High Performance Computing & Simulation (HPCS)

自引率

0.00%

发文量