Diagnosing Performance Bottlenecks in Massive Data Parallel Programs

2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid) Pub Date : 2016-05-16 DOI:10.1109/CCGrid.2016.81

Vinícius Dias, R. Moreira, Wagner Meira Jr, D. Guedes

引用次数: 7

Abstract

The increasing amount of data being stored and the variety of applications being proposed recently to make use of those data enabled a whole new generation of parallel programming environments and paradigms. Although most of these novel environments provide abstract programming interfaces and embed several run-time strategies that simplify several typical tasks in parallel and distributed systems, achieving good performance is still a challenge. In this paper we identify some common sources of performance degradation in the Spark programming environment and discuss some diagnosis dimensions that can be used to better understand such degradation. We then describe our experience in the use of those dimensions to drive the identification performance problems, and suggest how their impact may be minimized considering real applications.

查看原文本刊更多论文

海量数据并行程序的性能瓶颈诊断

存储的数据量的增加以及最近提出的利用这些数据的各种应用程序使新一代并行编程环境和范式成为可能。尽管这些新环境中的大多数都提供了抽象的编程接口，并嵌入了一些运行时策略，以简化并行和分布式系统中的一些典型任务，但实现良好的性能仍然是一个挑战。在本文中，我们确定了Spark编程环境中性能下降的一些常见来源，并讨论了一些可以用来更好地理解这种下降的诊断维度。然后，我们描述了我们在使用这些维度来驱动识别性能问题方面的经验，并建议如何考虑实际应用程序来最小化它们的影响。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid)

自引率

0.00%

发文量