使用多维视图表示数据集质量元数据

Joint Conference on Lexical and Computational Semantics Pub Date : 2014-08-11 DOI:10.1145/2660517.2660525

Jeremy Debattista, C. Lange, S. Auer

{"title":"使用多维视图表示数据集质量元数据","authors":"Jeremy Debattista, C. Lange, S. Auer","doi":"10.1145/2660517.2660525","DOIUrl":null,"url":null,"abstract":"Data quality is commonly defined as fitness for use. The problem of identifying quality of data is faced by many data consumers. Data publishers often do not have the means to identify quality problems in their data. To make the task for both stakeholders easier, we have developed the Dataset Quality Ontology (daQ). daQ is a core vocabulary for representing the results of quality benchmarking of a linked dataset. It represents quality metadata as multi-dimensional and statistical observations using the Data Cube vocabulary. Quality metadata are organised as a self-contained graph, which can, e.g., be embedded into linked open datasets. We discuss the design considerations, give examples for extending daQ by custom quality metrics, and present use cases such as analysing data versions, browsing datasets by quality, and link identification. We finally discuss how data cube visualisation tools enable data publishers and consumers to analyse better the quality of their data.","PeriodicalId":344435,"journal":{"name":"Joint Conference on Lexical and Computational Semantics","volume":"8 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2014-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"29","resultStr":"{\"title\":\"Representing dataset quality metadata using multi-dimensional views\",\"authors\":\"Jeremy Debattista, C. Lange, S. Auer\",\"doi\":\"10.1145/2660517.2660525\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Data quality is commonly defined as fitness for use. The problem of identifying quality of data is faced by many data consumers. Data publishers often do not have the means to identify quality problems in their data. To make the task for both stakeholders easier, we have developed the Dataset Quality Ontology (daQ). daQ is a core vocabulary for representing the results of quality benchmarking of a linked dataset. It represents quality metadata as multi-dimensional and statistical observations using the Data Cube vocabulary. Quality metadata are organised as a self-contained graph, which can, e.g., be embedded into linked open datasets. We discuss the design considerations, give examples for extending daQ by custom quality metrics, and present use cases such as analysing data versions, browsing datasets by quality, and link identification. We finally discuss how data cube visualisation tools enable data publishers and consumers to analyse better the quality of their data.\",\"PeriodicalId\":344435,\"journal\":{\"name\":\"Joint Conference on Lexical and Computational Semantics\",\"volume\":\"8 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2014-08-11\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"29\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Joint Conference on Lexical and Computational Semantics\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/2660517.2660525\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Joint Conference on Lexical and Computational Semantics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/2660517.2660525","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 29

摘要

数据质量通常被定义为适合使用。识别数据质量是许多数据使用者面临的问题。数据发布者通常没有办法识别其数据中的质量问题。为了使这两个利益相关者的任务更容易，我们开发了数据集质量本体(daQ)。daQ是表示链接数据集的质量基准测试结果的核心词汇。它使用Data Cube词汇表将高质量的元数据表示为多维和统计观察。高质量的元数据被组织成一个自包含的图，例如，可以嵌入到链接的开放数据集中。我们讨论了设计考虑因素，给出了通过自定义质量度量扩展daQ的示例，并给出了用例，例如分析数据版本、按质量浏览数据集和链接识别。我们最后讨论了数据立方体可视化工具如何使数据发布者和消费者能够更好地分析他们的数据质量。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Representing dataset quality metadata using multi-dimensional views

Data quality is commonly defined as fitness for use. The problem of identifying quality of data is faced by many data consumers. Data publishers often do not have the means to identify quality problems in their data. To make the task for both stakeholders easier, we have developed the Dataset Quality Ontology (daQ). daQ is a core vocabulary for representing the results of quality benchmarking of a linked dataset. It represents quality metadata as multi-dimensional and statistical observations using the Data Cube vocabulary. Quality metadata are organised as a self-contained graph, which can, e.g., be embedded into linked open datasets. We discuss the design considerations, give examples for extending daQ by custom quality metrics, and present use cases such as analysing data versions, browsing datasets by quality, and link identification. We finally discuss how data cube visualisation tools enable data publishers and consumers to analyse better the quality of their data.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Joint Conference on Lexical and Computational Semantics

自引率

0.00%

发文量