Reading between the lines of failure logs: Understanding how HPC systems fail

2013 43rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) Pub Date : 2013-06-24 DOI:10.1109/DSN.2013.6575356

Nosayba El-Sayed, Bianca Schroeder

引用次数: 103

Abstract

As the component count in supercomputing installations continues to increase, system reliability is becoming one of the major issues in designing HPC systems. These issues will become more challenging in future Exascale systems, which are predicted to include millions of CPU cores. Even with relatively reliable individual components, the sheer number of components will increase failure rates to unprecedented levels. Efficiently running those systems will require a good understanding of how different factors impact system reliability. In this paper we use a decade worth of field data made available by Los Alamos National Lab to study the impact of a diverse set of factors on the reliability of HPC systems. We provide insights into the nature of correlations between failures, and investigate the impact of factors, such as the power quality, temperature, fan and chiller reliability, system usage and utilization, and external factors, such as cosmic radiation, on system reliability.

查看原文本刊更多论文

阅读故障日志的字里行间:了解HPC系统是如何失败的

随着超级计算装置中组件数量的不断增加，系统可靠性成为设计高性能计算系统的主要问题之一。在未来的百亿亿级系统中，这些问题将变得更具挑战性，预计将包括数百万个CPU内核。即使使用相对可靠的单个组件，组件的数量也会将故障率提高到前所未有的水平。有效地运行这些系统需要很好地理解不同因素如何影响系统可靠性。在本文中，我们使用洛斯阿拉莫斯国家实验室提供的十年现场数据来研究各种因素对高性能计算系统可靠性的影响。我们提供对故障之间相关性本质的见解，并调查因素的影响，如电能质量，温度，风扇和冷却器可靠性，系统使用和利用率，以及外部因素，如宇宙辐射，对系统可靠性的影响。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2013 43rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN)

自引率

0.00%

发文量