RPViT:基于区域建议的视觉转换器

Proceedings of the 2022 5th International Conference on Image and Graphics Processing Pub Date : 2022-01-07 DOI:10.1145/3512388.3512421

Jing Ge, Qianxiang Wang, Jiahui Tong, Guangyu Gao

{"title":"RPViT:基于区域建议的视觉转换器","authors":"Jing Ge, Qianxiang Wang, Jiahui Tong, Guangyu Gao","doi":"10.1145/3512388.3512421","DOIUrl":null,"url":null,"abstract":"Vision Transformers constantly absorb the characteristics of convolutional neural networks to solve its shortcomings in translational invariance and scale invariance. However, dividing the image by a simple grid often destroys the position and scale features in the image at the beginning of the network. In this paper, we propose a vision transformer based on region proposal, which obtains the inductive bias in a simple way. Specifically, RPViT achieves locality and scale-invariance by extracting regions with locality using a traditional region proposal algorithm and deflating objects of different scales to the same scale by a bilinear interpolation algorithm. In addition, to enable the network to fully utilize and encode diverse candidate objects, a multi-class token approach based on orthogonalization is proposed and applied. Experiments on ImageNet demonstrate that RPViT outperforms baseline converters and related work.","PeriodicalId":434878,"journal":{"name":"Proceedings of the 2022 5th International Conference on Image and Graphics Processing","volume":"414 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-01-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"RPViT: Vision Transformer Based on Region Proposal\",\"authors\":\"Jing Ge, Qianxiang Wang, Jiahui Tong, Guangyu Gao\",\"doi\":\"10.1145/3512388.3512421\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Vision Transformers constantly absorb the characteristics of convolutional neural networks to solve its shortcomings in translational invariance and scale invariance. However, dividing the image by a simple grid often destroys the position and scale features in the image at the beginning of the network. In this paper, we propose a vision transformer based on region proposal, which obtains the inductive bias in a simple way. Specifically, RPViT achieves locality and scale-invariance by extracting regions with locality using a traditional region proposal algorithm and deflating objects of different scales to the same scale by a bilinear interpolation algorithm. In addition, to enable the network to fully utilize and encode diverse candidate objects, a multi-class token approach based on orthogonalization is proposed and applied. Experiments on ImageNet demonstrate that RPViT outperforms baseline converters and related work.\",\"PeriodicalId\":434878,\"journal\":{\"name\":\"Proceedings of the 2022 5th International Conference on Image and Graphics Processing\",\"volume\":\"414 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-01-07\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 2022 5th International Conference on Image and Graphics Processing\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/3512388.3512421\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2022 5th International Conference on Image and Graphics Processing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3512388.3512421","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

Vision transformer不断吸收卷积神经网络的特点，解决卷积神经网络在平移不变性和尺度不变性方面的不足。然而，用简单的网格划分图像往往会破坏网络开始时图像中的位置和尺度特征。本文提出了一种基于区域建议的视觉变压器，以一种简单的方法获得感应偏置。具体而言，RPViT通过传统的区域建议算法提取具有局部性的区域，并通过双线性插值算法将不同尺度的对象压缩到相同尺度，从而实现局部性和尺度不变性。此外，为了使网络能够充分利用和编码各种候选对象，提出并应用了一种基于正交化的多类令牌方法。在ImageNet上的实验表明，RPViT优于基准转换器和相关工作。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

RPViT: Vision Transformer Based on Region Proposal

Vision Transformers constantly absorb the characteristics of convolutional neural networks to solve its shortcomings in translational invariance and scale invariance. However, dividing the image by a simple grid often destroys the position and scale features in the image at the beginning of the network. In this paper, we propose a vision transformer based on region proposal, which obtains the inductive bias in a simple way. Specifically, RPViT achieves locality and scale-invariance by extracting regions with locality using a traditional region proposal algorithm and deflating objects of different scales to the same scale by a bilinear interpolation algorithm. In addition, to enable the network to fully utilize and encode diverse candidate objects, a multi-class token approach based on orthogonalization is proposed and applied. Experiments on ImageNet demonstrate that RPViT outperforms baseline converters and related work.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings of the 2022 5th International Conference on Image and Graphics Processing

自引率

0.00%

发文量