Revisiting the Double Checkpointing Algorithm

2013 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum Pub Date : 2013-05-20 DOI:10.1109/IPDPSW.2013.11

J. Dongarra, T. Hérault, Y. Robert

引用次数: 14

Abstract

Fast check pointing algorithms require distributed access to stable storage. This paper revisits the approach base upon double check pointing, and compares the blocking algorithm of Zheng, Shi and Kalé, with the non-blocking algorithm of Ni, Meneses and Kalé, in terms of both performance and risk. We also extend their model proposed to assess the impact of the overhead associated to non-blocking communications. We then provide a new peer-to-peer check pointing algorithm, called the triple check pointing algorithm, that can work at constant memory, and achieves both higher efficiency and better risk handling than the double check pointing algorithm. We provide performance and risk models for all the evaluated protocols, and compare them through comprehensive simulations.

查看原文本刊更多论文

重新审视双检查点算法

快速检查指向算法需要对稳定存储进行分布式访问。本文重新审视了基于双重检查点的方法，并将Zheng, Shi和kal的阻塞算法与Ni, Meneses和kal的非阻塞算法在性能和风险方面进行了比较。我们还扩展了他们提出的模型，以评估与非阻塞通信相关的开销的影响。然后，我们提供了一种新的点对点检查点算法，称为三重检查点算法，它可以在恒定内存下工作，并且比双重检查点算法具有更高的效率和更好的风险处理能力。我们提供了所有评估协议的性能和风险模型，并通过综合仿真对它们进行了比较。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2013 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum

自引率

0.00%

发文量