Endong Wang, Shaohua Wu, Qing Zhang, Jun Liu, Wenlu Zhang, Zhihong Lin, Yutong Lu, Yunfei Du, Xiaoqian Zhu
{"title":"The Gyrokinetic Particle Simulation of Fusion Plasmas on Tianhe-2 Supercomputer","authors":"Endong Wang, Shaohua Wu, Qing Zhang, Jun Liu, Wenlu Zhang, Zhihong Lin, Yutong Lu, Yunfei Du, Xiaoqian Zhu","doi":"10.1109/SCALA.2016.8","DOIUrl":null,"url":null,"abstract":"We present novel optimizations of the fusion plasmas simulation code, GTC on Tianhe-2 supercomputer. The simulation exhibits excellent weak scalability up to 3072 31S1P Xeon Phi co-processors. An unprecedented up to 5.8× performance improvement is achieved for the GTC on Tianhe-2. An efficient particle exchanging algorithm is developed that simplifies the original iterative scheme to a direct implementation, which leads to a 7.9× performance improvement in terms of MPI communications on 1024 nodes of Tianhe-2. A customized particle sorting algorithm is presented that delivers a 2.0× performance improvement on the co-processor for the kernel relating to the particle computing. A smart offload algorithm that minimizes the data exchange between host and co-processor is introduced. Other optimizations like the loop fusion and vectorization are also presented.","PeriodicalId":410521,"journal":{"name":"2016 7th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA)","volume":"64 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2016-11-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2016 7th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SCALA.2016.8","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 5
Abstract
We present novel optimizations of the fusion plasmas simulation code, GTC on Tianhe-2 supercomputer. The simulation exhibits excellent weak scalability up to 3072 31S1P Xeon Phi co-processors. An unprecedented up to 5.8× performance improvement is achieved for the GTC on Tianhe-2. An efficient particle exchanging algorithm is developed that simplifies the original iterative scheme to a direct implementation, which leads to a 7.9× performance improvement in terms of MPI communications on 1024 nodes of Tianhe-2. A customized particle sorting algorithm is presented that delivers a 2.0× performance improvement on the co-processor for the kernel relating to the particle computing. A smart offload algorithm that minimizes the data exchange between host and co-processor is introduced. Other optimizations like the loop fusion and vectorization are also presented.