Learning to Parallelize in a Shared-Memory Environment with Transformers

Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming Pub Date : 2022-04-27 DOI:10.1145/3572848.3582565

Re'em Harel, Yuval Pinter, Gal Oren

{"title":"Learning to Parallelize in a Shared-Memory Environment with Transformers","authors":"Re'em Harel, Yuval Pinter, Gal Oren","doi":"10.1145/3572848.3582565","DOIUrl":null,"url":null,"abstract":"In past years, the world has switched to multi and many core shared memory architectures. As a result, there is a growing need to utilize these architectures by introducing shared memory parallelization schemes, such as OpenMP, to applications. Nevertheless, introducing OpenMP work-sharing loop construct into code, especially legacy code, is challenging due to pervasive pitfalls in management of parallel shared memory. To facilitate the performance of this task, many source-to-source (S2S) compilers have been created over the years, tasked with inserting OpenMP directives into code automatically. In addition to having limited robustness to their input format, these compilers still do not achieve satisfactory coverage and precision in locating parallelizable code and generating appropriate directives. In this work, we propose leveraging recent advances in machine learning techniques, specifically in natural language processing (NLP), to suggest the need for an OpenMP work-sharing loop directive and data-sharing attributes clauses --- the building blocks of concurrent programming. We train several transformer models, named PragFormer, for these tasks and show that they outperform statistically-trained baselines and automatic source-to-source (S2S) parallelization compilers in both classifying the overall need for an parallel for directive and the introduction of private and reduction clauses. In the future, our corpus can be used for additional tasks, up to generating entire OpenMP directives. The source code and database for our project can be accessed on GitHub 1 and HuggingFace 2.","PeriodicalId":233744,"journal":{"name":"Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming","volume":"17 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-04-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"8","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3572848.3582565","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 8

Abstract

In past years, the world has switched to multi and many core shared memory architectures. As a result, there is a growing need to utilize these architectures by introducing shared memory parallelization schemes, such as OpenMP, to applications. Nevertheless, introducing OpenMP work-sharing loop construct into code, especially legacy code, is challenging due to pervasive pitfalls in management of parallel shared memory. To facilitate the performance of this task, many source-to-source (S2S) compilers have been created over the years, tasked with inserting OpenMP directives into code automatically. In addition to having limited robustness to their input format, these compilers still do not achieve satisfactory coverage and precision in locating parallelizable code and generating appropriate directives. In this work, we propose leveraging recent advances in machine learning techniques, specifically in natural language processing (NLP), to suggest the need for an OpenMP work-sharing loop directive and data-sharing attributes clauses --- the building blocks of concurrent programming. We train several transformer models, named PragFormer, for these tasks and show that they outperform statistically-trained baselines and automatic source-to-source (S2S) parallelization compilers in both classifying the overall need for an parallel for directive and the introduction of private and reduction clauses. In the future, our corpus can be used for additional tasks, up to generating entire OpenMP directives. The source code and database for our project can be accessed on GitHub 1 and HuggingFace 2.

查看原文本刊更多论文

学习使用变压器在共享内存环境中并行化

在过去的几年里，世界已经转向多核和多核共享内存架构。因此，越来越需要通过向应用程序引入共享内存并行化方案(如OpenMP)来利用这些体系结构。然而，由于并行共享内存管理中的普遍缺陷，将OpenMP工作共享循环结构引入代码(特别是遗留代码)是具有挑战性的。为了促进这项任务的执行，多年来已经创建了许多源到源(S2S)编译器，其任务是将OpenMP指令自动插入代码中。除了对其输入格式具有有限的健壮性之外，这些编译器在定位可并行代码和生成适当指令方面仍然没有达到令人满意的覆盖率和精度。在这项工作中，我们建议利用机器学习技术的最新进展，特别是在自然语言处理(NLP)方面，建议需要OpenMP工作共享循环指令和数据共享属性子句-并发编程的构建块。我们为这些任务训练了几个名为PragFormer的转换器模型，并表明它们在对并行指令的总体需求进行分类以及引入私有和缩减子句方面优于统计训练的基线和自动源到源(S2S)并行编译器。在未来，我们的语料库可以用于其他任务，直至生成整个OpenMP指令。我们项目的源代码和数据库可以在GitHub 1和HuggingFace 2上访问。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming

自引率

0.00%

发文量