Using checkpointing to recover from poor multi-site parallel job scheduling decisions

Middleware for Grid Computing Pub Date : 2007-11-26 DOI:10.1145/1376849.1376851

William M. Jones

引用次数: 5

Abstract

Recent research in multi-site parallel job scheduling leverages user-provided estimates of job communication characteristics to effectively partition the job across multiple clusters. Previous research addressed the impact of inaccuracies in these estimates on overall system performance and found that multi-site scheduling techniques benefit from these estimates, even in the presence of considerable inaccuracy. While these results are encouraging, there are many instances where these errors result in poor scheduling decisions that cause network over-subscription. This situation can lead to significantly degraded application runtime performance and turnaround time. In this paper, we explore the use of job checkpointing to selectively stop offending jobs in order to alleviate network congestion and subsequently restart them when (and where) sufficient network resources are available. We then characterize the conditions and the extent to which checkpointing improves overall performance. We demonstrate that checkpointing is beneficial even when the overhead of doing so is costly.

查看原文本刊更多论文

使用检查点从糟糕的多站点并行作业调度决策中恢复

最近的多站点并行作业调度研究利用用户提供的作业通信特征估计来有效地跨多个集群划分作业。先前的研究解决了这些估计的不准确性对整体系统性能的影响，并发现多站点调度技术受益于这些估计，即使在存在相当大的不准确性的情况下。虽然这些结果令人鼓舞，但在许多情况下，这些错误会导致糟糕的调度决策，从而导致网络过度订阅。这种情况可能导致应用程序运行时性能和周转时间显著降低。在本文中，我们探索了使用作业检查点来选择性地停止违规作业，以缓解网络拥塞，并随后在足够的网络资源可用时(以及在哪里)重新启动它们。然后，我们描述了检查点提高整体性能的条件和程度。我们证明了检查点是有益的，即使这样做的开销是昂贵的。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Middleware for Grid Computing

自引率

0.00%

发文量