Stellar Mergers with HPX-Kokkos and SYCL: Methods of using an Asynchronous Many-Task Runtime System with SYCL

Proceedings of the 2023 International Workshop on OpenCL Pub Date : 2023-03-04 DOI:10.1145/3585341.3585354

Gregor Daiß, Patrick Diehl, H. Kaiser, D. Pflüger

{"title":"Stellar Mergers with HPX-Kokkos and SYCL: Methods of using an Asynchronous Many-Task Runtime System with SYCL","authors":"Gregor Daiß, Patrick Diehl, H. Kaiser, D. Pflüger","doi":"10.1145/3585341.3585354","DOIUrl":null,"url":null,"abstract":"Ranging from NVIDIA GPUs to AMD GPUs and Intel GPUs: Given the heterogeneity of available accelerator cards within current supercomputers, portability is a key aspect for modern HPC applications. In Octo-Tiger, an astrophysics application simulating binary star systems and stellar mergers, we rely on Kokkos and its various execution spaces for portable compute kernels. In turn, we use HPX, a distributed task-based runtime system, to coordinate kernel launches, CPU tasks, and communication. This combination allows us to have a fine interleaving between portable CPU/GPU computations and communication, enabling scalability on various supercomputers. However, for HPX and Kokkos to work together optimally, we need to be able to treat Kokkos kernels as HPX tasks. Otherwise, instead of integrating asynchronous Kokkos kernel launches into HPX’s task graph, we would have to actively wait for them with fence commands, which wastes CPU time better spent otherwise. Using an integration layer called HPX-Kokkos, treating Kokkos kernels as tasks already works for some Kokkos execution spaces (like the CUDA one), but not for others (like the SYCL one). In this work, we started making Octo-Tiger and HPX itself compatible with SYCL. To do so, we introduce numerous software changes most notably an HPX-SYCL integration. This integration allows us to treat SYCL events as HPX tasks, which in turn allows us to better integrate Kokkos by extending the support of HPX-Kokkos to also fully support Kokkos’ SYCL execution space. We show two ways to implement this HPX-SYCL integration and test them using Octo-Tiger and its Kokkos kernels, on both an NVIDIA A100 and an AMD MI100. We find modest, yet noticeable, speedups (1.11x to 1.15x for the relevant configurations) by enabling this integration, even when just running simple single-node scenarios with Octo-Tiger where communication and CPU utilization are not yet an issue. We further find that the integration using event polling within the HPX scheduler works far better than the alternative implementation using SYCL host tasks.","PeriodicalId":360830,"journal":{"name":"Proceedings of the 2023 International Workshop on OpenCL","volume":"36 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-03-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2023 International Workshop on OpenCL","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3585341.3585354","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

Abstract

Ranging from NVIDIA GPUs to AMD GPUs and Intel GPUs: Given the heterogeneity of available accelerator cards within current supercomputers, portability is a key aspect for modern HPC applications. In Octo-Tiger, an astrophysics application simulating binary star systems and stellar mergers, we rely on Kokkos and its various execution spaces for portable compute kernels. In turn, we use HPX, a distributed task-based runtime system, to coordinate kernel launches, CPU tasks, and communication. This combination allows us to have a fine interleaving between portable CPU/GPU computations and communication, enabling scalability on various supercomputers. However, for HPX and Kokkos to work together optimally, we need to be able to treat Kokkos kernels as HPX tasks. Otherwise, instead of integrating asynchronous Kokkos kernel launches into HPX’s task graph, we would have to actively wait for them with fence commands, which wastes CPU time better spent otherwise. Using an integration layer called HPX-Kokkos, treating Kokkos kernels as tasks already works for some Kokkos execution spaces (like the CUDA one), but not for others (like the SYCL one). In this work, we started making Octo-Tiger and HPX itself compatible with SYCL. To do so, we introduce numerous software changes most notably an HPX-SYCL integration. This integration allows us to treat SYCL events as HPX tasks, which in turn allows us to better integrate Kokkos by extending the support of HPX-Kokkos to also fully support Kokkos’ SYCL execution space. We show two ways to implement this HPX-SYCL integration and test them using Octo-Tiger and its Kokkos kernels, on both an NVIDIA A100 and an AMD MI100. We find modest, yet noticeable, speedups (1.11x to 1.15x for the relevant configurations) by enabling this integration, even when just running simple single-node scenarios with Octo-Tiger where communication and CPU utilization are not yet an issue. We further find that the integration using event polling within the HPX scheduler works far better than the alternative implementation using SYCL host tasks.

查看原文本刊更多论文

恒星合并与hx - kokkos和SYCL:使用异步多任务运行时系统与SYCL的方法

从NVIDIA gpu到AMD gpu和Intel gpu:考虑到当前超级计算机中可用加速卡的异质性，可移植性是现代HPC应用程序的关键方面。在模拟双星系统和恒星合并的天体物理学应用程序Octo-Tiger中，我们依赖Kokkos及其各种可移植计算内核的执行空间。然后，我们使用HPX(一个基于分布式任务的运行时系统)来协调内核启动、CPU任务和通信。这种组合使我们能够在便携式CPU/GPU计算和通信之间实现良好的交叉，从而在各种超级计算机上实现可扩展性。然而，为了使HPX和Kokkos能够最佳地协同工作，我们需要能够将Kokkos内核视为HPX任务。否则，不是将异步Kokkos内核启动集成到HPX的任务图中，我们将不得不使用fence命令积极地等待它们，这会浪费CPU时间。使用名为HPX-Kokkos的集成层，将Kokkos内核作为任务处理已经适用于某些Kokkos执行空间(如CUDA)，但不适用于其他空间(如SYCL)。在这项工作中，我们开始使Octo-Tiger和HPX本身与SYCL兼容。为此，我们引入了许多软件更改，最显著的是HPX-SYCL集成。这种集成允许我们将SYCL事件视为HPX任务，这反过来又允许我们通过扩展HPX-Kokkos的支持来更好地集成Kokkos，从而也完全支持Kokkos的SYCL执行空间。我们展示了两种实现HPX-SYCL集成的方法，并在NVIDIA A100和AMD MI100上使用Octo-Tiger及其Kokkos内核进行了测试。通过启用这种集成，我们发现了适度但明显的速度提升(相关配置为1.11倍到1.15倍)，即使只是在使用Octo-Tiger运行简单的单节点场景时也是如此，其中通信和CPU利用率还不是问题。我们进一步发现，在HPX调度器中使用事件轮询的集成比使用SYCL主机任务的替代实现要好得多。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 2023 International Workshop on OpenCL

自引率

0.00%

发文量