利用自动生成的合成图像数据集对人脸识别进行基准测试

2021 IEEE International Joint Conference on Biometrics (IJCB) Pub Date : 2021-06-08 DOI:10.1109/IJCB52358.2021.9484363

Laurent Colbois, Tiago de Freitas Pereira, S. Marcel

{"title":"利用自动生成的合成图像数据集对人脸识别进行基准测试","authors":"Laurent Colbois, Tiago de Freitas Pereira, S. Marcel","doi":"10.1109/IJCB52358.2021.9484363","DOIUrl":null,"url":null,"abstract":"The availability of large-scale face datasets has been key in the progress of face recognition. However, due to licensing issues or copyright infringement, some datasets are not available anymore (e.g. MS-Celeb-1M). Recent advances in Generative Adversarial Networks (GANs), to synthesize realistic face images, provide a pathway to replace real datasets by synthetic datasets, both to train and benchmark face recognition (FR) systems. The work presented in this paper provides a study on benchmarking FR systems using a synthetic dataset. First, we introduce the proposed methodology to generate a synthetic dataset, without the need for human intervention, by exploiting the latent structure of a StyleGAN2 model with multiple controlled factors of variation. Then, we confirm that (i) the generated synthetic identities are not data subjects from the GAN’s training dataset, which is verified on a synthetic dataset with 10K+ identities; (ii) benchmarking results on the synthetic dataset are a good substitution, often providing error rates and system ranking similar to the benchmarking on the real dataset.","PeriodicalId":175984,"journal":{"name":"2021 IEEE International Joint Conference on Biometrics (IJCB)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-06-08","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"18","resultStr":"{\"title\":\"On the use of automatically generated synthetic image datasets for benchmarking face recognition\",\"authors\":\"Laurent Colbois, Tiago de Freitas Pereira, S. Marcel\",\"doi\":\"10.1109/IJCB52358.2021.9484363\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The availability of large-scale face datasets has been key in the progress of face recognition. However, due to licensing issues or copyright infringement, some datasets are not available anymore (e.g. MS-Celeb-1M). Recent advances in Generative Adversarial Networks (GANs), to synthesize realistic face images, provide a pathway to replace real datasets by synthetic datasets, both to train and benchmark face recognition (FR) systems. The work presented in this paper provides a study on benchmarking FR systems using a synthetic dataset. First, we introduce the proposed methodology to generate a synthetic dataset, without the need for human intervention, by exploiting the latent structure of a StyleGAN2 model with multiple controlled factors of variation. Then, we confirm that (i) the generated synthetic identities are not data subjects from the GAN’s training dataset, which is verified on a synthetic dataset with 10K+ identities; (ii) benchmarking results on the synthetic dataset are a good substitution, often providing error rates and system ranking similar to the benchmarking on the real dataset.\",\"PeriodicalId\":175984,\"journal\":{\"name\":\"2021 IEEE International Joint Conference on Biometrics (IJCB)\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-06-08\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"18\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2021 IEEE International Joint Conference on Biometrics (IJCB)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/IJCB52358.2021.9484363\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 IEEE International Joint Conference on Biometrics (IJCB)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IJCB52358.2021.9484363","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 18

摘要

大规模人脸数据集的可用性是人脸识别进展的关键。然而，由于许可问题或版权侵权，一些数据集不再可用(例如MS-Celeb-1M)。生成对抗网络(GANs)的最新进展，以合成逼真的人脸图像，提供了一种途径，用合成数据集取代真实数据集，用于训练和基准人脸识别(FR)系统。本文提出的工作提供了使用合成数据集对FR系统进行基准测试的研究。首先，我们介绍了所提出的方法，通过利用具有多个控制变量的StyleGAN2模型的潜在结构来生成合成数据集，而无需人工干预。然后，我们确认(i)生成的合成身份不是来自GAN训练数据集的数据主体，这在具有10K+身份的合成数据集上进行了验证;(ii)合成数据集上的基准测试结果是一个很好的替代，通常提供与真实数据集上的基准测试相似的错误率和系统排名。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

On the use of automatically generated synthetic image datasets for benchmarking face recognition

The availability of large-scale face datasets has been key in the progress of face recognition. However, due to licensing issues or copyright infringement, some datasets are not available anymore (e.g. MS-Celeb-1M). Recent advances in Generative Adversarial Networks (GANs), to synthesize realistic face images, provide a pathway to replace real datasets by synthetic datasets, both to train and benchmark face recognition (FR) systems. The work presented in this paper provides a study on benchmarking FR systems using a synthetic dataset. First, we introduce the proposed methodology to generate a synthetic dataset, without the need for human intervention, by exploiting the latent structure of a StyleGAN2 model with multiple controlled factors of variation. Then, we confirm that (i) the generated synthetic identities are not data subjects from the GAN’s training dataset, which is verified on a synthetic dataset with 10K+ identities; (ii) benchmarking results on the synthetic dataset are a good substitution, often providing error rates and system ranking similar to the benchmarking on the real dataset.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2021 IEEE International Joint Conference on Biometrics (IJCB)

自引率

0.00%

发文量