在全动态数据流上维护具有恢复功能的 $k$-MinHash 签名

arXiv - CS - Data Structures and Algorithms Pub Date : 2024-07-31 DOI:arxiv-2407.21614

Andrea Clementi, Luciano Gualà, Luca Pepè Sciarria, Alessandro Straziota

{"title":"在全动态数据流上维护具有恢复功能的 $k$-MinHash 签名","authors":"Andrea Clementi, Luciano Gualà, Luca Pepè Sciarria, Alessandro Straziota","doi":"arxiv-2407.21614","DOIUrl":null,"url":null,"abstract":"We consider the task of performing Jaccard similarity queries over a large\ncollection of items that are dynamically updated according to a streaming input\nmodel. An item here is a subset of a large universe $U$ of elements. A\nwell-studied approach to address this important problem in data mining is to\ndesign fast-similarity data sketches. In this paper, we focus on global\nsolutions for this problem, i.e., a single data structure which is able to\nanswer both Similarity Estimation and All-Candidate Pairs queries, while also\ndynamically managing an arbitrary, online sequence of element insertions and\ndeletions received in input. We introduce and provide an in-depth analysis of a dynamic, buffered version\nof the well-known $k$-MinHash sketch. This buffered version better manages\ncritical update operations thus significantly reducing the number of times the\nsketch needs to be rebuilt from scratch using expensive recovery queries. We\nprove that the buffered $k$-MinHash uses $O(k \\log |U|)$ memory words per\nsubset and that its amortized update time per insertion/deletion is $O(k \\log\n|U|)$ with high probability. Moreover, our data structure can return the\n$k$-MinHash signature of any subset in $O(k)$ time, and this signature is\nexactly the same signature that would be computed from scratch (and thus the\nquality of the signature is the same as the one guaranteed by the static\n$k$-MinHash). Analytical and experimental comparisons with the other, state-of-the-art\nglobal solutions for this problem given in [Bury et al.,WSDM'18] show that the\nbuffered $k$-MinHash turns out to be competitive in a wide and relevant range\nof the online input parameters.","PeriodicalId":501525,"journal":{"name":"arXiv - CS - Data Structures and Algorithms","volume":"209 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2024-07-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Maintaining $k$-MinHash Signatures over Fully-Dynamic Data Streams with Recovery\",\"authors\":\"Andrea Clementi, Luciano Gualà, Luca Pepè Sciarria, Alessandro Straziota\",\"doi\":\"arxiv-2407.21614\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"We consider the task of performing Jaccard similarity queries over a large\\ncollection of items that are dynamically updated according to a streaming input\\nmodel. An item here is a subset of a large universe $U$ of elements. A\\nwell-studied approach to address this important problem in data mining is to\\ndesign fast-similarity data sketches. In this paper, we focus on global\\nsolutions for this problem, i.e., a single data structure which is able to\\nanswer both Similarity Estimation and All-Candidate Pairs queries, while also\\ndynamically managing an arbitrary, online sequence of element insertions and\\ndeletions received in input. We introduce and provide an in-depth analysis of a dynamic, buffered version\\nof the well-known $k$-MinHash sketch. This buffered version better manages\\ncritical update operations thus significantly reducing the number of times the\\nsketch needs to be rebuilt from scratch using expensive recovery queries. We\\nprove that the buffered $k$-MinHash uses $O(k \\\\log |U|)$ memory words per\\nsubset and that its amortized update time per insertion/deletion is $O(k \\\\log\\n|U|)$ with high probability. Moreover, our data structure can return the\\n$k$-MinHash signature of any subset in $O(k)$ time, and this signature is\\nexactly the same signature that would be computed from scratch (and thus the\\nquality of the signature is the same as the one guaranteed by the static\\n$k$-MinHash). Analytical and experimental comparisons with the other, state-of-the-art\\nglobal solutions for this problem given in [Bury et al.,WSDM'18] show that the\\nbuffered $k$-MinHash turns out to be competitive in a wide and relevant range\\nof the online input parameters.\",\"PeriodicalId\":501525,\"journal\":{\"name\":\"arXiv - CS - Data Structures and Algorithms\",\"volume\":\"209 1\",\"pages\":\"\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2024-07-31\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"arXiv - CS - Data Structures and Algorithms\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/arxiv-2407.21614\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"arXiv - CS - Data Structures and Algorithms","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/arxiv-2407.21614","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

我们考虑的任务是对根据流式输入模型动态更新的大量项目集合执行 Jaccard 相似性查询。这里的条目是由大量元素组成的$U$宇宙的一个子集。为解决数据挖掘中的这一重要问题，一种经过深入研究的方法是设计快速相似性数据草图。在本文中，我们将重点研究这一问题的全局解决方案，即能够同时回答相似性估计和全候选对查询的单一数据结构，同时还能动态管理输入中收到的任意、在线元素插入和删除序列。我们介绍并深入分析了著名的 $k$-MinHash 草图的动态缓冲版本。这种缓冲版本能更好地管理关键更新操作，从而大大减少了使用昂贵的恢复查询从头开始重建草图的次数。我们证明，缓冲版的$k$-MinHash每个子集使用了$O(k \log |U|)$内存字，而且每次插入/删除的摊销更新时间很有可能是$O(k \log |U|)$。此外，我们的数据结构可以在 $O(k)$ 时间内返回任意子集的$k$-MinHash 签名，而且该签名与从头开始计算的签名完全相同（因此签名的质量与静态$k$-MinHash 保证的质量相同）。与[Bury等人，WSDM'18]中针对这个问题给出的其他最先进的全局解决方案进行的分析和实验比较表明，缓冲式$k$-MinHash在广泛的相关在线输入参数范围内都具有竞争力。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Maintaining $k$-MinHash Signatures over Fully-Dynamic Data Streams with Recovery

We consider the task of performing Jaccard similarity queries over a large collection of items that are dynamically updated according to a streaming input model. An item here is a subset of a large universe $U$ of elements. A well-studied approach to address this important problem in data mining is to design fast-similarity data sketches. In this paper, we focus on global solutions for this problem, i.e., a single data structure which is able to answer both Similarity Estimation and All-Candidate Pairs queries, while also dynamically managing an arbitrary, online sequence of element insertions and deletions received in input. We introduce and provide an in-depth analysis of a dynamic, buffered version of the well-known $k$-MinHash sketch. This buffered version better manages critical update operations thus significantly reducing the number of times the sketch needs to be rebuilt from scratch using expensive recovery queries. We prove that the buffered $k$-MinHash uses $O(k \log |U|)$ memory words per subset and that its amortized update time per insertion/deletion is $O(k \log |U|)$ with high probability. Moreover, our data structure can return the $k$-MinHash signature of any subset in $O(k)$ time, and this signature is exactly the same signature that would be computed from scratch (and thus the quality of the signature is the same as the one guaranteed by the static $k$-MinHash). Analytical and experimental comparisons with the other, state-of-the-art global solutions for this problem given in [Bury et al.,WSDM'18] show that the buffered $k$-MinHash turns out to be competitive in a wide and relevant range of the online input parameters.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

arXiv - CS - Data Structures and Algorithms

自引率

0.00%

发文量