深度学习负载下新型人工智能加速器的综合评价

2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS) Pub Date : 2022-11-01 DOI:10.1109/PMBS56514.2022.00007

M. Emani, Zhen Xie, Siddhisanket Raskar, V. Sastry, William Arnold, Bruce Wilson, R. Thakur, V. Vishwanath, Zhengchun Liu, M. Papka, Cindy Orozco Bohorquez, Rickey C. Weisner, K. Li, Yongning Sheng, Yun Du, Jian Zhang, A. Tsyplikhin, Gurdaman S. Khaira, J. Fowers, R. Sivakumar, Victoria Godsoe, Adrián Macías, Chetan Tekur, Matthew Boyd

{"title":"深度学习负载下新型人工智能加速器的综合评价","authors":"M. Emani, Zhen Xie, Siddhisanket Raskar, V. Sastry, William Arnold, Bruce Wilson, R. Thakur, V. Vishwanath, Zhengchun Liu, M. Papka, Cindy Orozco Bohorquez, Rickey C. Weisner, K. Li, Yongning Sheng, Yun Du, Jian Zhang, A. Tsyplikhin, Gurdaman S. Khaira, J. Fowers, R. Sivakumar, Victoria Godsoe, Adrián Macías, Chetan Tekur, Matthew Boyd","doi":"10.1109/PMBS56514.2022.00007","DOIUrl":null,"url":null,"abstract":"Scientific applications are increasingly adopting Artificial Intelligence (AI) techniques to advance science. High-performance computing centers are evaluating emerging novel hardware accelerators to efficiently run AI-driven science applications. With a wide diversity in the hardware architectures and software stacks of these systems, it is challenging to understand how these accelerators perform. The state-of-the-art in the evaluation of deep learning workloads primarily focuses on CPUs and GPUs. In this paper, we present an overview of dataflow-based novel AI accelerators from SambaNova, Cerebras, Graphcore, and Groq. We present a first-of-a-kind evaluation of these accelerators with diverse workloads, such as Deep Learning (DL) primitives, benchmark models, and scientific machine learning applications. We also evaluate the performance of collective communication, which is key for distributed DL implementation, along with a study of scaling efficiency. We then discuss key insights, challenges, and opportunities in integrating these novel AI accelerators in supercomputing systems.","PeriodicalId":321991,"journal":{"name":"2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"A Comprehensive Evaluation of Novel AI Accelerators for Deep Learning Workloads\",\"authors\":\"M. Emani, Zhen Xie, Siddhisanket Raskar, V. Sastry, William Arnold, Bruce Wilson, R. Thakur, V. Vishwanath, Zhengchun Liu, M. Papka, Cindy Orozco Bohorquez, Rickey C. Weisner, K. Li, Yongning Sheng, Yun Du, Jian Zhang, A. Tsyplikhin, Gurdaman S. Khaira, J. Fowers, R. Sivakumar, Victoria Godsoe, Adrián Macías, Chetan Tekur, Matthew Boyd\",\"doi\":\"10.1109/PMBS56514.2022.00007\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Scientific applications are increasingly adopting Artificial Intelligence (AI) techniques to advance science. High-performance computing centers are evaluating emerging novel hardware accelerators to efficiently run AI-driven science applications. With a wide diversity in the hardware architectures and software stacks of these systems, it is challenging to understand how these accelerators perform. The state-of-the-art in the evaluation of deep learning workloads primarily focuses on CPUs and GPUs. In this paper, we present an overview of dataflow-based novel AI accelerators from SambaNova, Cerebras, Graphcore, and Groq. We present a first-of-a-kind evaluation of these accelerators with diverse workloads, such as Deep Learning (DL) primitives, benchmark models, and scientific machine learning applications. We also evaluate the performance of collective communication, which is key for distributed DL implementation, along with a study of scaling efficiency. We then discuss key insights, challenges, and opportunities in integrating these novel AI accelerators in supercomputing systems.\",\"PeriodicalId\":321991,\"journal\":{\"name\":\"2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-11-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/PMBS56514.2022.00007\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/PMBS56514.2022.00007","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

摘要

科学应用越来越多地采用人工智能(AI)技术来推进科学。高性能计算中心正在评估新兴的新型硬件加速器，以有效运行人工智能驱动的科学应用程序。由于这些系统的硬件架构和软件堆栈的多样性，理解这些加速器的性能是具有挑战性的。深度学习工作负载评估的最新技术主要集中在cpu和gpu上。在本文中，我们介绍了SambaNova, Cerebras, Graphcore和Groq基于数据流的新型AI加速器的概述。我们对这些具有不同工作负载的加速器进行了首次评估，例如深度学习(DL)原语、基准模型和科学机器学习应用程序。我们还评估了集体通信的性能，这是分布式深度学习实现的关键，同时还研究了扩展效率。然后，我们讨论了将这些新型人工智能加速器集成到超级计算系统中的关键见解、挑战和机遇。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

A Comprehensive Evaluation of Novel AI Accelerators for Deep Learning Workloads

Scientific applications are increasingly adopting Artificial Intelligence (AI) techniques to advance science. High-performance computing centers are evaluating emerging novel hardware accelerators to efficiently run AI-driven science applications. With a wide diversity in the hardware architectures and software stacks of these systems, it is challenging to understand how these accelerators perform. The state-of-the-art in the evaluation of deep learning workloads primarily focuses on CPUs and GPUs. In this paper, we present an overview of dataflow-based novel AI accelerators from SambaNova, Cerebras, Graphcore, and Groq. We present a first-of-a-kind evaluation of these accelerators with diverse workloads, such as Deep Learning (DL) primitives, benchmark models, and scientific machine learning applications. We also evaluate the performance of collective communication, which is key for distributed DL implementation, along with a study of scaling efficiency. We then discuss key insights, challenges, and opportunities in integrating these novel AI accelerators in supercomputing systems.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2022 IEEE/ACM International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)

自引率

0.00%

发文量