Unsupervised Cross-Lingual Speech Emotion Recognition Using Domain Adversarial Neural Network

2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) Pub Date : 2020-12-21 DOI:10.1109/ISCSLP49672.2021.9362058

Xiong Cai, Zhiyong Wu, Kuo Zhong, Bin Su, Dongyang Dai, H. Meng

{"title":"Unsupervised Cross-Lingual Speech Emotion Recognition Using Domain Adversarial Neural Network","authors":"Xiong Cai, Zhiyong Wu, Kuo Zhong, Bin Su, Dongyang Dai, H. Meng","doi":"10.1109/ISCSLP49672.2021.9362058","DOIUrl":null,"url":null,"abstract":"By using deep learning approaches, Speech Emotion Recognition (SER) on a single domain has achieved many excellent results. However, cross-domain SER is still a challenging task due to the distribution shift between source and target domains. In this work, we propose a Domain Adversarial Neural Network (DANN) based approach to mitigate this distribution shift problem for cross-lingual SER. Specifically, we add a language classifier and gradient reversal layer after the feature extractor to force the learned representation both language-independent and emotion-meaningful. Our method is unsupervised, i. e., labels on target language are not required, which makes it easier to apply our method to other languages. Experimental results show the proposed method provides an average absolute improvement of 3.91% over the baseline system for arousal and valence classification task. Furthermore, we find that batch normalization is beneficial to the performance gain of DANN. Therefore we also explore the effect of different ways of data combination for batch normalization.","PeriodicalId":279828,"journal":{"name":"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)","volume":"17 2 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-12-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"9","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ISCSLP49672.2021.9362058","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 9

Abstract

By using deep learning approaches, Speech Emotion Recognition (SER) on a single domain has achieved many excellent results. However, cross-domain SER is still a challenging task due to the distribution shift between source and target domains. In this work, we propose a Domain Adversarial Neural Network (DANN) based approach to mitigate this distribution shift problem for cross-lingual SER. Specifically, we add a language classifier and gradient reversal layer after the feature extractor to force the learned representation both language-independent and emotion-meaningful. Our method is unsupervised, i. e., labels on target language are not required, which makes it easier to apply our method to other languages. Experimental results show the proposed method provides an average absolute improvement of 3.91% over the baseline system for arousal and valence classification task. Furthermore, we find that batch normalization is beneficial to the performance gain of DANN. Therefore we also explore the effect of different ways of data combination for batch normalization.

查看原文本刊更多论文

基于领域对抗神经网络的无监督跨语言语音情感识别

通过使用深度学习方法，语音情感识别(SER)在单个域上取得了许多优异的成绩。然而，由于源域和目标域之间的分布变化，跨域SER仍然是一项具有挑战性的任务。在这项工作中，我们提出了一种基于领域对抗神经网络(DANN)的方法来缓解跨语言SER的分布转移问题。具体来说，我们在特征提取器之后添加了语言分类器和梯度反转层，以迫使学习到的表示既独立于语言又具有情感意义。我们的方法是无监督的，即不需要在目标语言上标记，这使得我们的方法更容易应用于其他语言。实验结果表明，该方法对唤醒和价态分类任务的平均绝对效率比基线系统提高了3.91%。此外，我们发现批归一化有利于DANN的性能提升。因此，我们还探讨了不同的数据组合方式对批归一化的影响。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)

自引率

0.00%

发文量