Still Open Problems in Data Warehouse and Data Lake Research: extended abstract

2021 Eighth International Conference on Social Network Analysis, Management and Security (SNAMS) Pub Date : 2021-12-06 DOI:10.1109/SNAMS53716.2021.9732098

R. Wrembel

引用次数: 2

Abstract

During recent years, we observe a widespread of new data sources, especially all types of social media and IoT devices, which produce huge data volumes, whose content ranges from fully structured to totally unstructured. All these types of data are commonly referred to as big data. They are typically described by the three most important characteristics, called 3V [1], namely: an extremely large volume, a variety of data models and structures (data representations), as well as a high velocity at which data are generated. We argue that out of these three Vs, the most challenging is variety [2]. Such data need to be integrated and transformed into a common representation, which is suitable for analysis, in a similar manner as traditional (mainly table-like) data.

查看原文本刊更多论文

数据仓库与数据湖研究中尚待解决的问题:扩展摘要

近年来，我们观察到广泛的新数据源，特别是各种类型的社交媒体和物联网设备，产生了巨大的数据量，其内容从完全结构化到完全非结构化。所有这些类型的数据通常被称为大数据。它们通常由三个最重要的特征来描述，称为3V[1]，即:极大的体积，各种数据模型和结构(数据表示)，以及数据生成的高速度。我们认为，在这三个v中，最具挑战性的是多样性[2]。这些数据需要以与传统数据(主要是类似表格的数据)类似的方式集成并转换为适合于分析的通用表示。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2021 Eighth International Conference on Social Network Analysis, Management and Security (SNAMS)

自引率

0.00%

发文量