{"title":"ERR-Net: Facial Expression Removal and Recognition Network With Residual Image","authors":"Baishuang Li;Siyi Mo;Wenming Yang;Guijin Wang;Qingmin Liao","doi":"10.1109/TBIOM.2023.3250832","DOIUrl":null,"url":null,"abstract":"Facial expression recognition is an important part of computer vision and has attracted great attention. Although deep learning pushes forward the development of facial expression recognition, it still faces huge challenges due to unrelated factors such as identity, gender, and race. Inspired by decomposing an expression into two parts: neutral component and expression component, we define residual features and propose an end-to-end network framework named Expression Removal and Recognition Network (ERR-Net), which can simultaneously perform expression removal and recognition tasks. The residual features are represented in two ways: pixel level and facial landmark level. Our network focuses on interpreting the encoder’s output and corresponding its segments to expressions to maximize the inter-class distances. We explore the improved generative adversarial network to convert different expressions into neutral expressions (i.e., expression removal), take the residual images as the output, learn the expression components in the process, and realize the classification of expressions. Through sufficient ablation experiments, we have proved that various improvements added on the network have obvious effects. Experimental comparisons on two benchmarks CK+ and MMI demonstrate that our proposed ERR-Net surpasses the state-of-the-art methods in terms of accuracy.","PeriodicalId":73307,"journal":{"name":"IEEE transactions on biometrics, behavior, and identity science","volume":"5 4","pages":"425-434"},"PeriodicalIF":0.0000,"publicationDate":"2023-03-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE transactions on biometrics, behavior, and identity science","FirstCategoryId":"1085","ListUrlMain":"https://ieeexplore.ieee.org/document/10061589/","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
Facial expression recognition is an important part of computer vision and has attracted great attention. Although deep learning pushes forward the development of facial expression recognition, it still faces huge challenges due to unrelated factors such as identity, gender, and race. Inspired by decomposing an expression into two parts: neutral component and expression component, we define residual features and propose an end-to-end network framework named Expression Removal and Recognition Network (ERR-Net), which can simultaneously perform expression removal and recognition tasks. The residual features are represented in two ways: pixel level and facial landmark level. Our network focuses on interpreting the encoder’s output and corresponding its segments to expressions to maximize the inter-class distances. We explore the improved generative adversarial network to convert different expressions into neutral expressions (i.e., expression removal), take the residual images as the output, learn the expression components in the process, and realize the classification of expressions. Through sufficient ablation experiments, we have proved that various improvements added on the network have obvious effects. Experimental comparisons on two benchmarks CK+ and MMI demonstrate that our proposed ERR-Net surpasses the state-of-the-art methods in terms of accuracy.