Collating time-series resource data for system-wide job profiling

NOMS 2016 - 2016 IEEE/IFIP Network Operations and Management Symposium Pub Date : 2016-04-25 DOI:10.1109/NOMS.2016.7502958

V. Bumgardner, V. Marek, Ray L. Hyatt

引用次数: 4

Abstract

Through the collection and association of discrete time-series resource metrics and workloads, we can both provide benchmark and intra-job resource collations, along with system-wide job profiling. Traditional RDBMSes are not designed to store and process long-term discrete time-series metrics and the commonly used resolution-reducing round robin databases (RRDB), make poor long-term sources of data for workload analytics. We implemented a system that employs “Big-data” (Hadoop/HBase) and other analytics (R) techniques and tools to store, process, and characterize HPC workloads. Using this system we have collected and processed over a 30 billion time-series metrics from existing short-term high-resolution (15-sec RRDB) sources, profiling over 200 thousand jobs across a wide spectrum of workloads. The system is currently in use at the University of Kentucky for better understanding of individual jobs and system-wide profiling as well as a strategic source of data for resource allocation and future acquisitions.

查看原文本刊更多论文

整理时间序列资源数据，用于系统范围的作业分析

通过收集和关联离散时间序列资源指标和工作负载，我们可以提供基准和作业内部资源排序，以及系统范围的作业分析。传统的rdbms并不是为存储和处理长期离散时间序列指标而设计的，而常用的降低分辨率的轮询数据库(RRDB)不能作为工作负载分析的长期数据源。我们实现了一个使用“大数据”(Hadoop/HBase)和其他分析(R)技术和工具来存储、处理和表征HPC工作负载的系统。使用该系统，我们已经从现有的短期高分辨率(15秒RRDB)来源收集和处理了超过300亿个时间序列指标，分析了各种工作负载中的20多万个作业。该系统目前在肯塔基大学使用，用于更好地了解单个工作和系统范围的分析，以及资源分配和未来采购的战略数据来源。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

NOMS 2016 - 2016 IEEE/IFIP Network Operations and Management Symposium

自引率

0.00%

发文量