Numerical Defect Correction as an Algorithm-Based Fault Tolerance Technique for Iterative Solvers

2011 IEEE 17th Pacific Rim International Symposium on Dependable Computing Pub Date : 2011-07-01 DOI:10.1109/PRDC.2011.26

Fabian Oboril, M. Tahoori, V. Heuveline, D. Lukarski, Jan-Philipp Weiss

引用次数: 27

Abstract

As hardware devices like processor cores and memory sub-systems based on nano-scale technology nodes become more unreliable, the need for fault tolerant numerical computing engines, as used in many critical applications with long computation/mission times, is becoming pronounced. In this paper, we present an Algorithm-based Fault Tolerance (ABFT) scheme for an iterative linear solver engine based on the Conjugated Gradient method (CG) by taking the advantage of numerical defect correction. This method is "pay as you go", meaning that there is practically only a runtime overhead if errors occur and a correction is performed. Our experimental comparison with software-based Triple Modular Redundancy (TMR) clearly shows the runtime benefit of the proposed approach, good fault tolerance and no occurrence of silent data corruption.

查看原文本刊更多论文

数值缺陷校正作为一种基于算法的迭代求解容错技术

随着基于纳米级技术节点的处理器内核和内存子系统等硬件设备变得越来越不可靠，在许多计算/任务时间长的关键应用中使用的容错数值计算引擎的需求变得越来越明显。本文利用数值缺陷校正的优势，提出了一种基于共轭梯度法(CG)的迭代线性求解引擎的算法容错方案。此方法是“随用随付”，这意味着如果发生错误并执行更正，实际上只有运行时开销。我们与基于软件的三模冗余(TMR)的实验比较清楚地表明，该方法的运行时优势，良好的容错性和不发生无声数据损坏。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2011 IEEE 17th Pacific Rim International Symposium on Dependable Computing

自引率

0.00%

发文量