CATCH: A Cloud-Based Adaptive Data Transfer Service for HPC

2011 IEEE International Parallel & Distributed Processing Symposium Pub Date : 2011-05-16 DOI:10.1109/IPDPS.2011.118

H. M. Monti, A. Butt, Sudharshan S. Vazhkudai

{"title":"CATCH: A Cloud-Based Adaptive Data Transfer Service for HPC","authors":"H. M. Monti, A. Butt, Sudharshan S. Vazhkudai","doi":"10.1109/IPDPS.2011.118","DOIUrl":null,"url":null,"abstract":"Modern High Performance Computing (HPC) applications process very large amounts of data. A critical research challenge lies in transporting input data to the HPC center from a number of distributed sources, e.g., scientific experiments and web repositories, etc., and offloading the result data to geographically distributed, intermittently available end-users, often over under-provisioned connections. Such end-user data services are typically performed using point-to-point transfers that are designed for well-endowed sites and are unable to reconcile the center's resource usage and users' delivery deadlines, unable to adapt to changing dynamics in the end-to-end data path and are not fault-tolerant. To overcome these inefficiencies, decentralized HPC data services are emerging as viable alternatives. In this paper, we develop and enhance such distributed data services by designing CATCH, a Cloud-based Adaptive data Transfer service for HPC. CATCH leverages a bevy of cloud storage resources to orchestrate a decentralized data transport with fail-over capabilities. Our results demonstrate that CATCH is a feasible approach, and can help improve the data transfer times at the HPC center by as much as 81.1\\% for typical HPC workloads.","PeriodicalId":355100,"journal":{"name":"2011 IEEE International Parallel & Distributed Processing Symposium","volume":"63 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2011-05-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"32","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2011 IEEE International Parallel & Distributed Processing Symposium","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IPDPS.2011.118","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 32

Abstract

Modern High Performance Computing (HPC) applications process very large amounts of data. A critical research challenge lies in transporting input data to the HPC center from a number of distributed sources, e.g., scientific experiments and web repositories, etc., and offloading the result data to geographically distributed, intermittently available end-users, often over under-provisioned connections. Such end-user data services are typically performed using point-to-point transfers that are designed for well-endowed sites and are unable to reconcile the center's resource usage and users' delivery deadlines, unable to adapt to changing dynamics in the end-to-end data path and are not fault-tolerant. To overcome these inefficiencies, decentralized HPC data services are emerging as viable alternatives. In this paper, we develop and enhance such distributed data services by designing CATCH, a Cloud-based Adaptive data Transfer service for HPC. CATCH leverages a bevy of cloud storage resources to orchestrate a decentralized data transport with fail-over capabilities. Our results demonstrate that CATCH is a feasible approach, and can help improve the data transfer times at the HPC center by as much as 81.1\% for typical HPC workloads.

查看原文本刊更多论文

捕获:基于云的HPC自适应数据传输服务

现代高性能计算(HPC)应用程序处理非常大量的数据。一个关键的研究挑战在于将输入数据从许多分布式来源(例如，科学实验和web存储库等)传输到HPC中心，并将结果数据卸载给地理上分布的、间歇性可用的最终用户，通常是在供应不足的连接上。这种终端用户数据服务通常使用点对点传输来执行，这种传输是为资源丰富的站点设计的，无法协调中心的资源使用和用户的交付期限，无法适应端到端数据路径中不断变化的动态，并且不具有容错性。为了克服这些低效率，分散的HPC数据服务正在成为可行的替代方案。本文通过设计基于云的HPC自适应数据传输服务CATCH来开发和增强这种分布式数据服务。CATCH利用一组云存储资源来编排具有故障转移功能的分散数据传输。我们的研究结果表明，CATCH是一种可行的方法，对于典型的HPC工作负载，它可以帮助将HPC中心的数据传输时间提高81.1%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2011 IEEE International Parallel & Distributed Processing Symposium

自引率

0.00%

发文量