Dongju Cha , Jaewook Lee , Daeyoung Jung, Sangheon Pack
{"title":"Fast and fair split computing for accelerating deep neural network (DNN) inference","authors":"Dongju Cha , Jaewook Lee , Daeyoung Jung, Sangheon Pack","doi":"10.1016/j.icte.2024.09.013","DOIUrl":null,"url":null,"abstract":"<div><div>Conventional split computing approaches for AI models that generate large outputs suffer from long transmission and inference times. Due to the limited resources of the edge server and selfish MDs, some MDs cannot offload their tasks and sacrifice their performance. To address these issues, we formulate an optimization problem to determine one or two split points that minimize inference latency while ensuring fair offloading among MDs. Additionally, we devise a low-complexity heuristic algorithm called fast and fair split computing (F2SC). Evaluation results demonstrate that F2SC reduces inference time by <span><math><mrow><mn>3</mn><mo>.</mo><mn>8</mn><mtext>%</mtext><mo>∼</mo><mn>20</mn><mo>.</mo><mn>1</mn><mtext>%</mtext></mrow></math></span> compared to the conventional approaches while maintaining fairness.</div></div>","PeriodicalId":48526,"journal":{"name":"ICT Express","volume":"11 1","pages":"Pages 47-52"},"PeriodicalIF":4.1000,"publicationDate":"2025-02-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ICT Express","FirstCategoryId":"94","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2405959524001164","RegionNum":3,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, INFORMATION SYSTEMS","Score":null,"Total":0}
引用次数: 0
Abstract
Conventional split computing approaches for AI models that generate large outputs suffer from long transmission and inference times. Due to the limited resources of the edge server and selfish MDs, some MDs cannot offload their tasks and sacrifice their performance. To address these issues, we formulate an optimization problem to determine one or two split points that minimize inference latency while ensuring fair offloading among MDs. Additionally, we devise a low-complexity heuristic algorithm called fast and fair split computing (F2SC). Evaluation results demonstrate that F2SC reduces inference time by compared to the conventional approaches while maintaining fairness.
期刊介绍:
The ICT Express journal published by the Korean Institute of Communications and Information Sciences (KICS) is an international, peer-reviewed research publication covering all aspects of information and communication technology. The journal aims to publish research that helps advance the theoretical and practical understanding of ICT convergence, platform technologies, communication networks, and device technologies. The technology advancement in information and communication technology (ICT) sector enables portable devices to be always connected while supporting high data rate, resulting in the recent popularity of smartphones that have a considerable impact in economic and social development.