Experimental evaluation of the fail-silent behaviour in programs with consistency checks

Proceedings of Annual Symposium on Fault Tolerant Computing Pub Date : 1996-06-25 DOI:10.1109/FTCS.1996.534625

M. Z. Rela, H. Madeira, J. G. Silva

{"title":"Experimental evaluation of the fail-silent behaviour in programs with consistency checks","authors":"M. Z. Rela, H. Madeira, J. G. Silva","doi":"10.1109/FTCS.1996.534625","DOIUrl":null,"url":null,"abstract":"An important research topic deals with the investigation of whether a non-duplicated computer can be made fail-silent, since that behaviour is a-priori assumed in many algorithms. However, previous research has shown that in systems using a simple behaviour based error detection mechanism invisible to the programmer (e.g. memory protection), the percentage of fail-silent violations could be higher than 10%. Since the study of these errors has shown that they were mostly caused by pure data errors, we evaluate the effectiveness of software techniques capable of checking the semantics of the data, such as assertions, to detect these remaining errors. The results of injecting physical pin-level faults show that these tests can prevent about 40% of the fail-silent model violations that escape the simple hardware-based error detection techniques. In order to decouple the intrinsic limitations of the tests used from other factors that might affect its error detection capabilities, we evaluated a special class of software checks known for its high theoretical coverage: algorithm based fault tolerance (ABFT). The analysis of the remaining errors showed that most of them remained undetected due to short range control flow errors. When very simple software-based control flow checking was associated to the semantic tests, the target system, without any dedicated error detection hardware, behaved according to the fail-silent model for about 98% of all the faults injected.","PeriodicalId":191163,"journal":{"name":"Proceedings of Annual Symposium on Fault Tolerant Computing","volume":"63 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1996-06-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"76","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of Annual Symposium on Fault Tolerant Computing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/FTCS.1996.534625","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 76

Abstract

An important research topic deals with the investigation of whether a non-duplicated computer can be made fail-silent, since that behaviour is a-priori assumed in many algorithms. However, previous research has shown that in systems using a simple behaviour based error detection mechanism invisible to the programmer (e.g. memory protection), the percentage of fail-silent violations could be higher than 10%. Since the study of these errors has shown that they were mostly caused by pure data errors, we evaluate the effectiveness of software techniques capable of checking the semantics of the data, such as assertions, to detect these remaining errors. The results of injecting physical pin-level faults show that these tests can prevent about 40% of the fail-silent model violations that escape the simple hardware-based error detection techniques. In order to decouple the intrinsic limitations of the tests used from other factors that might affect its error detection capabilities, we evaluated a special class of software checks known for its high theoretical coverage: algorithm based fault tolerance (ABFT). The analysis of the remaining errors showed that most of them remained undetected due to short range control flow errors. When very simple software-based control flow checking was associated to the semantic tests, the target system, without any dedicated error detection hardware, behaved according to the fail-silent model for about 98% of all the faults injected.

查看原文本刊更多论文

具有一致性检查的程序失效沉默行为的实验评价

一个重要的研究课题是调查一台非复制的计算机是否可以使故障沉默，因为这种行为在许多算法中是先验假设的。然而，先前的研究表明，在使用程序员看不见的基于简单行为的错误检测机制(例如内存保护)的系统中，故障静默违反的百分比可能高于10%。由于对这些错误的研究表明，它们主要是由纯粹的数据错误引起的，因此我们评估了能够检查数据语义(如断言)以检测这些剩余错误的软件技术的有效性。注入物理引脚级故障的结果表明，这些测试可以防止大约40%的故障沉默模型违规，这些违规无法通过简单的基于硬件的错误检测技术进行检测。为了将所使用的测试的内在限制与可能影响其错误检测能力的其他因素解耦，我们评估了一类特殊的软件检查，以其高理论覆盖率而闻名:基于算法的容错(ABFT)。对剩余误差的分析表明，由于控制流量误差较短，大多数误差未被检测到。当非常简单的基于软件的控制流检查与语义测试相关联时，目标系统在没有任何专用错误检测硬件的情况下，对大约98%的注入故障都按照故障沉默模型进行处理。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of Annual Symposium on Fault Tolerant Computing

自引率

0.00%

发文量