Speech emotion classification using semi-supervised LSTM

Advances in computational intelligence Pub Date : 2023-06-22 DOI:10.1007/s43674-023-00059-x

Nattipon Itponjaroen, Kumpee Apsornpasakorn, Eakarat Pimthai, Khwanchai Kaewkaisorn, Shularp Panitchart, Thitirat Siriborvornratanakul

{"title":"Speech emotion classification using semi-supervised LSTM","authors":"Nattipon Itponjaroen, Kumpee Apsornpasakorn, Eakarat Pimthai, Khwanchai Kaewkaisorn, Shularp Panitchart, Thitirat Siriborvornratanakul","doi":"10.1007/s43674-023-00059-x","DOIUrl":null,"url":null,"abstract":"<div><p>Speech mood analysis is a challenging task with unclear optimal feature selection. The nature of the dataset, whether it is from an infant or adult, is crucial to consider. In this study, the characteristics of speech were investigated using Mel-frequency cepstral coefficients (MFCC) to analyze audio files. The CREMA-D dataset, which includes six different mood states (normal, angry, happy, sad, scared, and irritated), was employed to identify mood states from speech files. A mood classification system was proposed that integrates Support Vector Machines (SVM) and Long Short-Term Memory (LSTM) models to increase the number of labeled data in small datasets and improve classification accuracy.</p><p>A semi-supervised model was proposed in this study to improve the accuracy of speech mood classification systems. The approach was tested on a classification model that used SVM and LSTM, and it was found that the semi-supervised model outperforms both SVM and LSTM models, achieving a validation accuracy of 89.72%. This result surpasses the accuracy achieved by SVM and LSTM models alone. Moreover, the semi-supervised method was observed to accelerate the training process of the model. These outcomes illustrate the efficacy of the proposed model and its potential to enhance speech mood analysis techniques.</p></div>","PeriodicalId":72089,"journal":{"name":"Advances in computational intelligence","volume":"3 4","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2023-06-22","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Advances in computational intelligence","FirstCategoryId":"1085","ListUrlMain":"https://link.springer.com/article/10.1007/s43674-023-00059-x","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Speech mood analysis is a challenging task with unclear optimal feature selection. The nature of the dataset, whether it is from an infant or adult, is crucial to consider. In this study, the characteristics of speech were investigated using Mel-frequency cepstral coefficients (MFCC) to analyze audio files. The CREMA-D dataset, which includes six different mood states (normal, angry, happy, sad, scared, and irritated), was employed to identify mood states from speech files. A mood classification system was proposed that integrates Support Vector Machines (SVM) and Long Short-Term Memory (LSTM) models to increase the number of labeled data in small datasets and improve classification accuracy.

A semi-supervised model was proposed in this study to improve the accuracy of speech mood classification systems. The approach was tested on a classification model that used SVM and LSTM, and it was found that the semi-supervised model outperforms both SVM and LSTM models, achieving a validation accuracy of 89.72%. This result surpasses the accuracy achieved by SVM and LSTM models alone. Moreover, the semi-supervised method was observed to accelerate the training process of the model. These outcomes illustrate the efficacy of the proposed model and its potential to enhance speech mood analysis techniques.

查看原文本刊更多论文

基于半监督LSTM的语音情感分类

语音情绪分析是一项具有挑战性的任务，最优特征选择不明确。数据集的性质，无论是来自婴儿还是成人，都是至关重要的。在本研究中，使用梅尔频率倒谱系数（MFCC）来分析音频文件，以研究语音的特征。CREMA-D数据集包括六种不同的情绪状态（正常、愤怒、快乐、悲伤、害怕和愤怒），用于从语音文件中识别情绪状态。提出了一种结合支持向量机（SVM）和长短期记忆（LSTM）模型的情绪分类系统，以增加小数据集中的标记数据数量，提高分类精度。为了提高语音情绪分类系统的准确性，本文提出了一种半监督模型。该方法在一个使用SVM和LSTM的分类模型上进行了测试，发现半监督模型优于SVM和LSTM模型，验证准确率达到89.72%。这一结果超过了单独使用SVM和LSTM模型的准确率。此外，观察到半监督方法加速了模型的训练过程。这些结果说明了所提出的模型的有效性及其增强语音情绪分析技术的潜力。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Advances in computational intelligence

自引率

0.00%

发文量