How to recover efficiently and asynchronously when optimism fails

Proceedings of 16th International Conference on Distributed Computing Systems Pub Date : 1996-05-27 DOI:10.1109/ICDCS.1996.507907

O. Damani, V. Garg

引用次数: 74

Abstract

We propose a new algorithm for recovering asynchronously from failures in a distributed computation. Our algorithm is based on two novel concepts-a fault-tolerant vector clock to maintain causality information in spite of failures, and a history mechanism to detect orphan states and obsolete messages. These two mechanisms together with checkpointing and message-logging are used to restore the system to a consistent state after a failure of one or more processes. Our algorithm is completely asynchronous. It handles multiple failures, does not assume any message ordering, causes the minimum amount of rollback and restores the maximum recoverable state with low overhead. Earlier optimistic protocols lack one or more of the above properties.

查看原文本刊更多论文

当乐观情绪失败时，如何高效异步地恢复

提出了一种分布式计算中异步恢复故障的新算法。我们的算法基于两个新概念——一个容错矢量时钟，用于在故障情况下保持因果关系信息;一个历史机制，用于检测孤立状态和过时消息。这两种机制与检查点和消息日志一起用于在一个或多个进程失败后将系统恢复到一致状态。我们的算法完全异步。它处理多个故障，不假定任何消息排序，导致最小数量的回滚，并以低开销恢复最大可恢复状态。早期的乐观协议缺少上述一个或多个属性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of 16th International Conference on Distributed Computing Systems

自引率

0.00%

发文量