{"title":"基于改进卷积神经网络的高效并行处理","authors":"Sang-Soo Park, Jung-Hyun Hong, Ki-Seok Chung","doi":"10.1109/IRI.2017.37","DOIUrl":null,"url":null,"abstract":"Today, Convolutional Neural Network (CNN) is adopted in a lot of areas such as computer vision and natural language processing. By employing hardware accelerators such as graphic processing unit (GPU), a significant amount of speedup can be achieved in CNN and many studies have proposed such acceleration methods. However, it is not straightforward to parallelize the CNN on a hardware accelerator because there are irregular characteristics of generating output feature maps. In this paper, we propose a modified CNN for efficient parallel processing. A well-known CNN architecture called Lenet-5 has an inefficient convolution combination. The proposed method of this paper improves the efficiency by utilizing a special operation called dummy operation. The proposed method is capable of maximizing the utilization of GPU by modifying Lenet-5's convolution combination. Its improved efficiency is validated on a platform that integrates a CPU and a GPU in the same die. Our OpenCL implementation of the proposed method has achieved an average peak performance of 115.66 GFLOPS which is an improvement of 37.26 times in execution time. Further, a reduction of 26.40 times in energy consumption is achieved.","PeriodicalId":254330,"journal":{"name":"2017 IEEE International Conference on Information Reuse and Integration (IRI)","volume":"9 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2017-08-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"Modified Convolution Neural Network for Highly Effective Parallel Processing\",\"authors\":\"Sang-Soo Park, Jung-Hyun Hong, Ki-Seok Chung\",\"doi\":\"10.1109/IRI.2017.37\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Today, Convolutional Neural Network (CNN) is adopted in a lot of areas such as computer vision and natural language processing. By employing hardware accelerators such as graphic processing unit (GPU), a significant amount of speedup can be achieved in CNN and many studies have proposed such acceleration methods. However, it is not straightforward to parallelize the CNN on a hardware accelerator because there are irregular characteristics of generating output feature maps. In this paper, we propose a modified CNN for efficient parallel processing. A well-known CNN architecture called Lenet-5 has an inefficient convolution combination. The proposed method of this paper improves the efficiency by utilizing a special operation called dummy operation. The proposed method is capable of maximizing the utilization of GPU by modifying Lenet-5's convolution combination. Its improved efficiency is validated on a platform that integrates a CPU and a GPU in the same die. Our OpenCL implementation of the proposed method has achieved an average peak performance of 115.66 GFLOPS which is an improvement of 37.26 times in execution time. Further, a reduction of 26.40 times in energy consumption is achieved.\",\"PeriodicalId\":254330,\"journal\":{\"name\":\"2017 IEEE International Conference on Information Reuse and Integration (IRI)\",\"volume\":\"9 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2017-08-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2017 IEEE International Conference on Information Reuse and Integration (IRI)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/IRI.2017.37\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2017 IEEE International Conference on Information Reuse and Integration (IRI)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IRI.2017.37","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Modified Convolution Neural Network for Highly Effective Parallel Processing
Today, Convolutional Neural Network (CNN) is adopted in a lot of areas such as computer vision and natural language processing. By employing hardware accelerators such as graphic processing unit (GPU), a significant amount of speedup can be achieved in CNN and many studies have proposed such acceleration methods. However, it is not straightforward to parallelize the CNN on a hardware accelerator because there are irregular characteristics of generating output feature maps. In this paper, we propose a modified CNN for efficient parallel processing. A well-known CNN architecture called Lenet-5 has an inefficient convolution combination. The proposed method of this paper improves the efficiency by utilizing a special operation called dummy operation. The proposed method is capable of maximizing the utilization of GPU by modifying Lenet-5's convolution combination. Its improved efficiency is validated on a platform that integrates a CPU and a GPU in the same die. Our OpenCL implementation of the proposed method has achieved an average peak performance of 115.66 GFLOPS which is an improvement of 37.26 times in execution time. Further, a reduction of 26.40 times in energy consumption is achieved.