Attribute Manipulation Generative Adversarial Networks for Fashion Images

2019 IEEE/CVF International Conference on Computer Vision (ICCV) Pub Date : 2019-10-01 DOI:10.1109/ICCV.2019.01064

Kenan E. Ak, A. Kassim, Joo-Hwee Lim, J. Y. Tham

{"title":"Attribute Manipulation Generative Adversarial Networks for Fashion Images","authors":"Kenan E. Ak, A. Kassim, Joo-Hwee Lim, J. Y. Tham","doi":"10.1109/ICCV.2019.01064","DOIUrl":null,"url":null,"abstract":"Recent advances in Generative Adversarial Networks (GANs) have made it possible to conduct multi-domain image-to-image translation using a single generative network. While recent methods such as Ganimation and SaGAN are able to conduct translations on attribute-relevant regions using attention, they do not perform well when the number of attributes increases as the training of attention masks mostly rely on classification losses. To address this and other limitations, we introduce Attribute Manipulation Generative Adversarial Networks (AMGAN) for fashion images. While AMGAN's generator network uses class activation maps (CAMs) to empower its attention mechanism, it also exploits perceptual losses by assigning reference (target) images based on attribute similarities. AMGAN incorporates an additional discriminator network that focuses on attribute-relevant regions to detect unrealistic translations. Additionally, AMGAN can be controlled to perform attribute manipulations on specific regions such as the sleeve or torso regions. Experiments show that AMGAN outperforms state-of-the-art methods using traditional evaluation metrics as well as an alternative one that is based on image retrieval.","PeriodicalId":6728,"journal":{"name":"2019 IEEE/CVF International Conference on Computer Vision (ICCV)","volume":"24 1","pages":"10540-10549"},"PeriodicalIF":0.0000,"publicationDate":"2019-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"69","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 IEEE/CVF International Conference on Computer Vision (ICCV)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICCV.2019.01064","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 69

Abstract

Recent advances in Generative Adversarial Networks (GANs) have made it possible to conduct multi-domain image-to-image translation using a single generative network. While recent methods such as Ganimation and SaGAN are able to conduct translations on attribute-relevant regions using attention, they do not perform well when the number of attributes increases as the training of attention masks mostly rely on classification losses. To address this and other limitations, we introduce Attribute Manipulation Generative Adversarial Networks (AMGAN) for fashion images. While AMGAN's generator network uses class activation maps (CAMs) to empower its attention mechanism, it also exploits perceptual losses by assigning reference (target) images based on attribute similarities. AMGAN incorporates an additional discriminator network that focuses on attribute-relevant regions to detect unrealistic translations. Additionally, AMGAN can be controlled to perform attribute manipulations on specific regions such as the sleeve or torso regions. Experiments show that AMGAN outperforms state-of-the-art methods using traditional evaluation metrics as well as an alternative one that is based on image retrieval.

查看原文本刊更多论文

时尚图像属性操作生成对抗网络

生成对抗网络(GANs)的最新进展使得使用单个生成网络进行多域图像到图像的翻译成为可能。虽然最近的方法，如Ganimation和SaGAN能够使用注意力在属性相关区域上进行翻译，但当属性数量增加时，它们表现不佳，因为注意力掩模的训练主要依赖于分类损失。为了解决这个问题和其他限制，我们为时尚图像引入了属性操作生成对抗网络(AMGAN)。虽然AMGAN的生成器网络使用类激活图(CAMs)来增强其注意机制，但它也通过基于属性相似性分配参考(目标)图像来利用感知损失。AMGAN结合了一个额外的判别器网络，该网络专注于属性相关区域，以检测不现实的翻译。此外，可以控制AMGAN对特定区域(如袖子或躯干区域)执行属性操作。实验表明，AMGAN优于使用传统评估指标的最先进方法以及基于图像检索的替代方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2019 IEEE/CVF International Conference on Computer Vision (ICCV)

自引率

0.00%

发文量