ACM SIGMOD Record最新文献

Auto-Tables: Relationalize Tables without Using Examples 自动表格无需使用示例即可建立关系表

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665269

Peng Li, Yeye He, Cong Yan, Yue Wang, Surajit Chaudhuri

引用次数: 0

From Binary Join to Free Join 从二进制加盟到免费加盟

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665259

Y. Wang, Max Willsey, Dan Suciu

引用次数: 0

Technical Perspective: Efficient and Reusable Lazy Sampling 技术视角：高效、可重复使用的懒惰采样

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665260

Thomas Neumann

引用次数: 0

Efficient and Reusable Lazy Sampling 高效、可重复使用的懒人取样

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665261

Viktor Sanca, Periklis Chrysogelos, Anastasia Ailamaki

{"title":"Efficient and Reusable Lazy Sampling","authors":"Viktor Sanca, Periklis Chrysogelos, Anastasia Ailamaki","doi":"10.1145/3665252.3665261","DOIUrl":"https://doi.org/10.1145/3665252.3665261","url":null,"abstract":"Modern analytical engines rely on Approximate Query Processing (AQP) to provide faster response times than the hardware allows for exact query answering. However, existing AQP methods impose steep performance penalties as workload unpredictability increases. While offline AQP relies on predictable workloads to a priori create samples that match the queries, as soon as workload predictability diminishes, returning to existing online AQP methods that create query-specific samples with little reuse across queries results in significantly smaller gains in response times. As a result, existing approaches cannot fully exploit the benefits of sampling under increased unpredictability.\u0000 We propose LAQy, a framework for building, expanding, and merging samples to adapt to the changes in workload predicates. We propose lazy sampling to overcome the unpredictability issues that cause fast-but-specialized samples to be query-specific and design it for a scale-up analytical engine to show the adaptivity and practicality of our framework in a modern system. LAQy speeds up online sampling processing as a function of data access and computation reuse, making sampler placement after expensive operators more practical.","PeriodicalId":346332,"journal":{"name":"ACM SIGMOD Record","volume":"27 7","pages":""},"PeriodicalIF":0.0,"publicationDate":"2024-05-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140979380","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

DBSP: Incremental Computation on Streams and Its Applications to Databases DBSP：流上的增量计算及其在数据库中的应用

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665271

Mihai Budiu, Tej Chajed, Frank McSherry, Leonid Ryzhyk, V. Tannen

引用次数: 0

Technical Perspective: Synthetic Data Needs a Reproducibility Benchmark 技术视角：合成数据需要可重复性基准

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665266

Xi He

引用次数: 0

Technical Perspective on 'Better Differentially Private Approximate Histograms and Heavy Hitters using the Misra-Gries Sketch' 关于 "使用米斯拉-格里斯草图获得更好的差分私有近似直方图和重击 "的技术视角

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665254

Graham Cormode

引用次数: 0

Unicorn: A Unified Multi-Tasking Matching Model 独角兽统一的多任务匹配模型

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665263

Ju Fan, Jianhong Tu, Guoliang Li, Peng Wang, Xiaoyong Du, Xiaofeng Jia, Song Gao, Nan Tang

{"title":"Unicorn: A Unified Multi-Tasking Matching Model","authors":"Ju Fan, Jianhong Tu, Guoliang Li, Peng Wang, Xiaoyong Du, Xiaofeng Jia, Song Gao, Nan Tang","doi":"10.1145/3665252.3665263","DOIUrl":"https://doi.org/10.1145/3665252.3665263","url":null,"abstract":"Data matching, which decides whether two data elements (e.g., string, tuple, column, or knowledge graph entity) are the \"same\" (a.k.a. a match), is a key concept in data integration. The widely used practice is to build task-specific or even dataset-specific solutions, which are hard to generalize and disable the opportunities of knowledge sharing that can be learned from different datasets and multiple tasks. In this paper, we propose Unicorn, a unified model for generally supporting common data matching tasks. Building such a unified model is challenging due to heterogeneous formats of input data elements and various matching semantics of multiple tasks. To address the challenges, Unicorn employs one generic Encoder that converts any pair of data elements (a, b) into a learned representation, and uses a Matcher, which is a binary classifier, to decide whether a matches b. To align matching semantics of multiple tasks, Unicorn adopts a mixture-of-experts model that enhances the learned representation into a better representation. We conduct extensive experiments using 20 datasets on 7 well-studied data matching tasks, and find that our unified model can achieve better performance on most tasks and on average, compared with the state-of-the-art specific models trained for ad-hoc tasks and datasets separately. Moreover, Unicorn can also well serve new matching tasks with zero-shot learning.","PeriodicalId":346332,"journal":{"name":"ACM SIGMOD Record","volume":"76 4","pages":""},"PeriodicalIF":0.0,"publicationDate":"2024-05-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140978783","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Technical Perspective: Graph Theory for Data Privacy: A New Approach for Complex Data Flows 技术视角：数据隐私的图论：复杂数据流的新方法

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665264

Elena Ferrari

{"title":"Technical Perspective: Graph Theory for Data Privacy: A New Approach for Complex Data Flows","authors":"Elena Ferrari","doi":"10.1145/3665252.3665264","DOIUrl":"https://doi.org/10.1145/3665252.3665264","url":null,"abstract":"Nearly all of the world's population now uses online services that request personal information, covering almost every aspect of our lives. The abundance of personal data in digital form has brought incredible benefits to end users, enabling them to access personalized and advanced services based on the analysis of the data collected. This capability has dramatically improved the user experience in various application domains, ranging from healthcare to e-commerce, finance, logistics, and entertainment, to name a few. Numerous technological advancements in the field of big data have enabled this massive processing of personal data, and recent advances in AI data processing capabilities will expand the ways in which service providers will use personal data in the coming years. Machine learning algorithms, powered by AI, will be used to make increasingly accurate predictions about user behavior by uncovering hidden correlations within massive data sets. There is therefore a tension between the desire to fully exploit personal data in such ecosystems and the need to provide strong privacy and transparency guarantees to the individuals whose data is being exploited. Privacy protection is further complicated because data processing is typically not performed in isolation but through pipelines of different services, with each step making inferences about the personal data consumed by the services in subsequent steps.","PeriodicalId":346332,"journal":{"name":"ACM SIGMOD Record","volume":"35 11","pages":""},"PeriodicalIF":0.0,"publicationDate":"2024-05-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140980059","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Learning to Restructure Tables Automatically 学会自动重组表格

ACM SIGMOD Record Pub Date : 2024-05-14 DOI: 10.1145/3665252.3665268

J. M. Hellerstein

引用次数: 0