越南语语音识别的端到端模型

Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies Pub Date : 2019-03-20 DOI:10.1109/RIVF.2019.8713758

V. Nguyen

{"title":"越南语语音识别的端到端模型","authors":"V. Nguyen","doi":"10.1109/RIVF.2019.8713758","DOIUrl":null,"url":null,"abstract":"This paper presents an approach of End-to-End model based on Long Short-Term Memory (LSTM) and Time Delay Deep Neural Network (TDNN) models for Vietnamese speech recognition. Two Vietnamese End-to-End architectures using Connectionist Temporal Classification (CTC) as the loss function are proposed. The paper also presents the method to construct phonesets based on Vietnamese characters or tonemes to produce the label sequence for any given transcription when applying CTC model. The experimental results showed that CTC based End-To-End models are competitive to traditional models with only 5% of WER worse for Vietnamese speech recognition (SR), but the advantage is no requirement of forced alignment for training acoustic models. In addition, it is similar to proposed studies using traditional models, the tone information and the toneme set are solutions to optimize the performance for Vietnamese SR.","PeriodicalId":171525,"journal":{"name":"Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies","volume":"71 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-03-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"4","resultStr":"{\"title\":\"An End-to-End Model for Vietnamese Speech Recognition\",\"authors\":\"V. Nguyen\",\"doi\":\"10.1109/RIVF.2019.8713758\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper presents an approach of End-to-End model based on Long Short-Term Memory (LSTM) and Time Delay Deep Neural Network (TDNN) models for Vietnamese speech recognition. Two Vietnamese End-to-End architectures using Connectionist Temporal Classification (CTC) as the loss function are proposed. The paper also presents the method to construct phonesets based on Vietnamese characters or tonemes to produce the label sequence for any given transcription when applying CTC model. The experimental results showed that CTC based End-To-End models are competitive to traditional models with only 5% of WER worse for Vietnamese speech recognition (SR), but the advantage is no requirement of forced alignment for training acoustic models. In addition, it is similar to proposed studies using traditional models, the tone information and the toneme set are solutions to optimize the performance for Vietnamese SR.\",\"PeriodicalId\":171525,\"journal\":{\"name\":\"Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies\",\"volume\":\"71 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-03-20\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"4\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/RIVF.2019.8713758\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/RIVF.2019.8713758","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 4

摘要

提出了一种基于长短期记忆(LSTM)和时延深度神经网络(TDNN)模型的端到端越南语语音识别方法。提出了两种使用连接时间分类(CTC)作为损失函数的越南端到端结构。本文还介绍了在使用CTC模型时，基于越南语字符或音素构建电话集以产生任意给定转录的标签序列的方法。实验结果表明，基于CTC的端到端模型在越南语语音识别(SR)中与传统模型相比仅差5%，但其优势在于训练声学模型时不需要强制对齐。此外，与使用传统模型的研究类似，声调信息和声调集是优化越南语SR性能的解决方案。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

An End-to-End Model for Vietnamese Speech Recognition

This paper presents an approach of End-to-End model based on Long Short-Term Memory (LSTM) and Time Delay Deep Neural Network (TDNN) models for Vietnamese speech recognition. Two Vietnamese End-to-End architectures using Connectionist Temporal Classification (CTC) as the loss function are proposed. The paper also presents the method to construct phonesets based on Vietnamese characters or tonemes to produce the label sequence for any given transcription when applying CTC model. The experimental results showed that CTC based End-To-End models are competitive to traditional models with only 5% of WER worse for Vietnamese speech recognition (SR), but the advantage is no requirement of forced alignment for training acoustic models. In addition, it is similar to proposed studies using traditional models, the tone information and the toneme set are solutions to optimize the performance for Vietnamese SR.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Conference on Research, Innovation and Vision for the Future in Computing & Communication Technologies

自引率

0.00%

发文量