Adaptive-precision SIMD architecture for high-throughput and resource-efficient DNN acceleration

IF 2.6 3区 工程技术 Q3 COMPUTER SCIENCE, HARDWARE & ARCHITECTURE
Integration-The Vlsi Journal Pub Date : 2026-05-01 Epub Date: 2026-01-16 DOI:10.1016/j.vlsi.2026.102666
Vasundhara Trivedi , Harman Singh Bagga , Gopal Raut , Santosh Kumar Vishvakarma
{"title":"Adaptive-precision SIMD architecture for high-throughput and resource-efficient DNN acceleration","authors":"Vasundhara Trivedi ,&nbsp;Harman Singh Bagga ,&nbsp;Gopal Raut ,&nbsp;Santosh Kumar Vishvakarma","doi":"10.1016/j.vlsi.2026.102666","DOIUrl":null,"url":null,"abstract":"<div><div>Deep Neural Network (DNN) accelerators require high computational throughput and flexible precision support while operating under stringent resource and power constraints. To address these challenges, we propose an adaptive-precision SIMD (Single Instruction, Multiple Data) Processing Element (PE) architecture for signed integer and fixed-point operations that maximizes resource utilization and enhances parallelism in multiply–accumulate (MAC) computations. The design introduces efficient resource reuse during partial product accumulation and supports both symmetric and asymmetric precision modes. Unlike state-of-the-art approaches, the proposed PE dynamically scales computation: processing 16 operands at low precision (4-bit), four operands at medium precision (8-bit), and a single operand at high precision (16-bit). Additionally, it supports asymmetric operations such as 16 <span><math><mo>×</mo></math></span> 4-bit multiplications in parallel, enabling unique flexibility and performance gains. The architecture is implemented and tested on ASIC and FPGA platforms. Accuracy evaluations across different DNN models and datasets show very small losses at reduced precision—less than 1% for LeNet on MNIST, 2.9% for AlexNet on CIFAR-10, 2.2% for VGG16 on CIFAR-10, and 3.5% for VGG16 on ImageNet-1000 compared to float32. Hardware synthesis yields significant improvements, including 46.2% fewer LUTs and 2.45 <span><math><mo>×</mo></math></span> less power on FPGA compared to existing designs. The proposed architecture delivers 2<span><math><mo>×</mo></math></span> higher throughput, upto 4.8<span><math><mo>×</mo></math></span> energy efficiency with 28.57% less area at 65 nm, compared to existing works, making it ideal for applications with variable precision and limited resources.</div></div>","PeriodicalId":54973,"journal":{"name":"Integration-The Vlsi Journal","volume":"108 ","pages":"Article 102666"},"PeriodicalIF":2.6000,"publicationDate":"2026-05-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Integration-The Vlsi Journal","FirstCategoryId":"5","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0167926026000210","RegionNum":3,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2026/1/16 0:00:00","PubModel":"Epub","JCR":"Q3","JCRName":"COMPUTER SCIENCE, HARDWARE & ARCHITECTURE","Score":null,"Total":0}
引用次数: 0

Abstract

Deep Neural Network (DNN) accelerators require high computational throughput and flexible precision support while operating under stringent resource and power constraints. To address these challenges, we propose an adaptive-precision SIMD (Single Instruction, Multiple Data) Processing Element (PE) architecture for signed integer and fixed-point operations that maximizes resource utilization and enhances parallelism in multiply–accumulate (MAC) computations. The design introduces efficient resource reuse during partial product accumulation and supports both symmetric and asymmetric precision modes. Unlike state-of-the-art approaches, the proposed PE dynamically scales computation: processing 16 operands at low precision (4-bit), four operands at medium precision (8-bit), and a single operand at high precision (16-bit). Additionally, it supports asymmetric operations such as 16 × 4-bit multiplications in parallel, enabling unique flexibility and performance gains. The architecture is implemented and tested on ASIC and FPGA platforms. Accuracy evaluations across different DNN models and datasets show very small losses at reduced precision—less than 1% for LeNet on MNIST, 2.9% for AlexNet on CIFAR-10, 2.2% for VGG16 on CIFAR-10, and 3.5% for VGG16 on ImageNet-1000 compared to float32. Hardware synthesis yields significant improvements, including 46.2% fewer LUTs and 2.45 × less power on FPGA compared to existing designs. The proposed architecture delivers 2× higher throughput, upto 4.8× energy efficiency with 28.57% less area at 65 nm, compared to existing works, making it ideal for applications with variable precision and limited resources.
用于高吞吐量和资源高效DNN加速的自适应精度SIMD架构
深度神经网络(DNN)加速器需要高计算吞吐量和灵活的精度支持,同时在严格的资源和功率限制下运行。为了解决这些挑战,我们提出了一种自适应精度SIMD(单指令,多数据)处理元素(PE)架构,用于有符号整数和定点操作,最大限度地提高资源利用率并增强乘法累积(MAC)计算的并行性。该设计在部分产品积累过程中引入了有效的资源重用,并支持对称和非对称精度模式。与最先进的方法不同,所提出的PE动态扩展计算:以低精度(4位)处理16个操作数,以中等精度(8位)处理4个操作数,以及以高精度(16位)处理单个操作数。此外,它还支持并行16 × 4位乘法等非对称操作,从而实现了独特的灵活性和性能提升。该体系结构在ASIC和FPGA平台上进行了实现和测试。不同DNN模型和数据集的精度评估显示,与float32相比,LeNet在MNIST上的精度损失非常小,AlexNet在CIFAR-10上的精度损失不到1%,VGG16在CIFAR-10上的精度损失不到2.2%,VGG16在ImageNet-1000上的精度损失不到3.5%。硬件合成产生了显著的改进,包括与现有设计相比,FPGA的lut减少了46.2%,功耗降低了2.45倍。与现有的架构相比,该架构的吞吐量提高了2倍,能效提高了4.8倍,65nm的面积减少了28.57%,使其成为可变精度和有限资源应用的理想选择。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 求助全文
来源期刊
Integration-The Vlsi Journal
Integration-The Vlsi Journal 工程技术-工程:电子与电气
CiteScore
3.80
自引率
5.30%
发文量
107
审稿时长
6 months
期刊介绍: Integration''s aim is to cover every aspect of the VLSI area, with an emphasis on cross-fertilization between various fields of science, and the design, verification, test and applications of integrated circuits and systems, as well as closely related topics in process and device technologies. Individual issues will feature peer-reviewed tutorials and articles as well as reviews of recent publications. The intended coverage of the journal can be assessed by examining the following (non-exclusive) list of topics: Specification methods and languages; Analog/Digital Integrated Circuits and Systems; VLSI architectures; Algorithms, methods and tools for modeling, simulation, synthesis and verification of integrated circuits and systems of any complexity; Embedded systems; High-level synthesis for VLSI systems; Logic synthesis and finite automata; Testing, design-for-test and test generation algorithms; Physical design; Formal verification; Algorithms implemented in VLSI systems; Systems engineering; Heterogeneous systems.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术官方微信
小红书