Transforming Image Super-Resolution: A ConvFormer-Based Efficient Approach

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society Pub Date : 2024-10-18 DOI:10.1109/TIP.2024.3477350

Gang Wu;Junjun Jiang;Junpeng Jiang;Xianming Liu

{"title":"Transforming Image Super-Resolution: A ConvFormer-Based Efficient Approach","authors":"Gang Wu;Junjun Jiang;Junpeng Jiang;Xianming Liu","doi":"10.1109/TIP.2024.3477350","DOIUrl":null,"url":null,"abstract":"Recent progress in single-image super-resolution (SISR) has achieved remarkable performance, yet the computational costs of these methods remain a challenge for deployment on resource-constrained devices. In particular, transformer-based methods, which leverage self-attention mechanisms, have led to significant breakthroughs but also introduce substantial computational costs. To tackle this issue, we introduce the Convolutional Transformer layer (ConvFormer) and propose a ConvFormer-based Super-Resolution network (CFSR), offering an effective and efficient solution for lightweight image super-resolution. The proposed method inherits the advantages of both convolution-based and transformer-based approaches. Specifically, CFSR utilizes large kernel convolutions as a feature mixer to replace the self-attention module, efficiently modeling long-range dependencies and extensive receptive fields with minimal computational overhead. Furthermore, we propose an edge-preserving feed-forward network (EFN) designed to achieve local feature aggregation while effectively preserving high-frequency information. Extensive experiments demonstrate that CFSR strikes an optimal balance between computational cost and performance compared to existing lightweight SR methods. When benchmarked against state-of-the-art methods such as ShuffleMixer, the proposed CFSR achieves a gain of 0.39 dB on the Urban100 dataset for the x2 super-resolution task while requiring 26% and 31% fewer parameters and FLOPs, respectively. The code and pre-trained models are available at \n<uri>https://github.com/Aitical/CFSR</uri>\n.","PeriodicalId":94032,"journal":{"name":"IEEE transactions on image processing : a publication of the IEEE Signal Processing Society","volume":"33 ","pages":"6071-6082"},"PeriodicalIF":0.0000,"publicationDate":"2024-10-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE transactions on image processing : a publication of the IEEE Signal Processing Society","FirstCategoryId":"1085","ListUrlMain":"https://ieeexplore.ieee.org/document/10723228/","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Recent progress in single-image super-resolution (SISR) has achieved remarkable performance, yet the computational costs of these methods remain a challenge for deployment on resource-constrained devices. In particular, transformer-based methods, which leverage self-attention mechanisms, have led to significant breakthroughs but also introduce substantial computational costs. To tackle this issue, we introduce the Convolutional Transformer layer (ConvFormer) and propose a ConvFormer-based Super-Resolution network (CFSR), offering an effective and efficient solution for lightweight image super-resolution. The proposed method inherits the advantages of both convolution-based and transformer-based approaches. Specifically, CFSR utilizes large kernel convolutions as a feature mixer to replace the self-attention module, efficiently modeling long-range dependencies and extensive receptive fields with minimal computational overhead. Furthermore, we propose an edge-preserving feed-forward network (EFN) designed to achieve local feature aggregation while effectively preserving high-frequency information. Extensive experiments demonstrate that CFSR strikes an optimal balance between computational cost and performance compared to existing lightweight SR methods. When benchmarked against state-of-the-art methods such as ShuffleMixer, the proposed CFSR achieves a gain of 0.39 dB on the Urban100 dataset for the x2 super-resolution task while requiring 26% and 31% fewer parameters and FLOPs, respectively. The code and pre-trained models are available at https://github.com/Aitical/CFSR .

查看原文本刊更多论文

转换图像超分辨率：基于 ConvFormer 的高效方法

单图像超分辨率（SISR）的最新进展取得了显著的性能，但这些方法的计算成本仍然是在资源受限的设备上部署所面临的挑战。特别是基于变压器的方法，这种方法利用了自注意机制，取得了重大突破，但也带来了巨大的计算成本。为了解决这个问题，我们引入了卷积变换器层（ConvFormer），并提出了基于卷积变换器的超分辨率网络（CFSR），为轻量级图像超分辨率提供了一种有效且高效的解决方案。所提出的方法继承了基于卷积和基于变换器两种方法的优点。具体来说，CFSR 利用大核卷积作为特征混合器来取代自注意模块，以最小的计算开销有效地模拟长程依赖性和广泛的感受野。此外，我们还提出了一种边缘保留前馈网络（EFN），旨在实现局部特征聚合，同时有效保留高频信息。大量实验证明，与现有的轻量级 SR 方法相比，CFSR 在计算成本和性能之间实现了最佳平衡。与 ShuffleMixer 等最先进的方法相比，所提出的 CFSR 在 Urban100 数据集上的 x2 超分辨率任务中实现了 0.39 dB 的增益，同时所需的参数和 FLOP 分别减少了 26% 和 31%。代码和预训练模型可在 https://github.com/Aitical/CFSR 上获取。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

IEEE transactions on image processing : a publication of the IEEE Signal Processing Society

自引率

0.00%

发文量