{"title":"基于高斯混合模型和KL散度的愤怒语和愤怒语分析","authors":"Shubham Mittal, Swati Vyas, S. Prasanna","doi":"10.1109/NCC.2013.6487985","DOIUrl":null,"url":null,"abstract":"Recognition of expressions from speech has emerged as an important research area in the recent past. However, the scientific community still faces problems in differentiating between angry and lombard speech. The objective of this work is to analyze the differences between the Lombard and angry speech using the features representing the excitation source of speech production. The instantaneous fundamental frequency, the strength of excitation and loudness measure, reflecting the sharpness of the impulse-like excitation around the epochs are used as excitation source features. The distributions curves of these three parameters are next plotted. We employ the concept of Gaussian Mixture Models (GMMs) and KL divergence (a measure of relative entropy) to calculate an exact measure of difference between angry, lombard and neutral speech with context to the aforementioned parameters and successfully show differences among the Lombard and angry speech signals at the excitation source level.","PeriodicalId":202526,"journal":{"name":"2013 National Conference on Communications (NCC)","volume":"11 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2013-03-28","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":"{\"title\":\"Analysis of lombard and angry speech using Gaussian Mixture Models and KL divergence\",\"authors\":\"Shubham Mittal, Swati Vyas, S. Prasanna\",\"doi\":\"10.1109/NCC.2013.6487985\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Recognition of expressions from speech has emerged as an important research area in the recent past. However, the scientific community still faces problems in differentiating between angry and lombard speech. The objective of this work is to analyze the differences between the Lombard and angry speech using the features representing the excitation source of speech production. The instantaneous fundamental frequency, the strength of excitation and loudness measure, reflecting the sharpness of the impulse-like excitation around the epochs are used as excitation source features. The distributions curves of these three parameters are next plotted. We employ the concept of Gaussian Mixture Models (GMMs) and KL divergence (a measure of relative entropy) to calculate an exact measure of difference between angry, lombard and neutral speech with context to the aforementioned parameters and successfully show differences among the Lombard and angry speech signals at the excitation source level.\",\"PeriodicalId\":202526,\"journal\":{\"name\":\"2013 National Conference on Communications (NCC)\",\"volume\":\"11 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2013-03-28\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"3\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2013 National Conference on Communications (NCC)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/NCC.2013.6487985\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2013 National Conference on Communications (NCC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/NCC.2013.6487985","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Analysis of lombard and angry speech using Gaussian Mixture Models and KL divergence
Recognition of expressions from speech has emerged as an important research area in the recent past. However, the scientific community still faces problems in differentiating between angry and lombard speech. The objective of this work is to analyze the differences between the Lombard and angry speech using the features representing the excitation source of speech production. The instantaneous fundamental frequency, the strength of excitation and loudness measure, reflecting the sharpness of the impulse-like excitation around the epochs are used as excitation source features. The distributions curves of these three parameters are next plotted. We employ the concept of Gaussian Mixture Models (GMMs) and KL divergence (a measure of relative entropy) to calculate an exact measure of difference between angry, lombard and neutral speech with context to the aforementioned parameters and successfully show differences among the Lombard and angry speech signals at the excitation source level.