基于变换的端到端语音识别中的位置编码研究

2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP) Pub Date : 2021-01-24 DOI:10.1109/ISCSLP49672.2021.9362093

Fengpeng Yue, Tom Ko

{"title":"基于变换的端到端语音识别中的位置编码研究","authors":"Fengpeng Yue, Tom Ko","doi":"10.1109/ISCSLP49672.2021.9362093","DOIUrl":null,"url":null,"abstract":"In the Transformer architecture, the model does not intrinsically learn the ordering information of the input frames and tokens due to its self-attention mechanism. In sequence-to-sequence learning tasks, the missing of ordering information is explicitly filled up by the use of positional representation. Currently, there are two major ways of using positional representation: the absolute way and relative way. In both ways, the positional in-formation is represented by positional vector. In this paper, we propose the use of positional matrix in the context of relative positional vector. Instead of adding the vectors to the key vectors in the self-attention layer, our method transforms the key vectors according to its position. Experiments on LibriSpeech dataset show that our approach outperforms the positional vector approach.","PeriodicalId":279828,"journal":{"name":"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)","volume":"140 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-01-24","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"An Investigation of Positional Encoding in Transformer-based End-to-end Speech Recognition\",\"authors\":\"Fengpeng Yue, Tom Ko\",\"doi\":\"10.1109/ISCSLP49672.2021.9362093\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In the Transformer architecture, the model does not intrinsically learn the ordering information of the input frames and tokens due to its self-attention mechanism. In sequence-to-sequence learning tasks, the missing of ordering information is explicitly filled up by the use of positional representation. Currently, there are two major ways of using positional representation: the absolute way and relative way. In both ways, the positional in-formation is represented by positional vector. In this paper, we propose the use of positional matrix in the context of relative positional vector. Instead of adding the vectors to the key vectors in the self-attention layer, our method transforms the key vectors according to its position. Experiments on LibriSpeech dataset show that our approach outperforms the positional vector approach.\",\"PeriodicalId\":279828,\"journal\":{\"name\":\"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)\",\"volume\":\"140 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2021-01-24\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ISCSLP49672.2021.9362093\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ISCSLP49672.2021.9362093","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 2

摘要

在Transformer体系结构中，由于其自关注机制，模型本质上并不学习输入帧和令牌的排序信息。在序列到序列的学习任务中，顺序信息的缺失通过使用位置表示来显式地填补。目前，位置表示的使用主要有两种方式:绝对方式和相对方式。在这两种方法中，位置信息都用位置向量表示。本文提出了位置矩阵在相对位置向量中的应用。我们的方法不是将向量添加到自关注层的关键向量上，而是根据关键向量的位置对其进行变换。在librisspeech数据集上的实验表明，该方法优于位置向量方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

An Investigation of Positional Encoding in Transformer-based End-to-end Speech Recognition

In the Transformer architecture, the model does not intrinsically learn the ordering information of the input frames and tokens due to its self-attention mechanism. In sequence-to-sequence learning tasks, the missing of ordering information is explicitly filled up by the use of positional representation. Currently, there are two major ways of using positional representation: the absolute way and relative way. In both ways, the positional in-formation is represented by positional vector. In this paper, we propose the use of positional matrix in the context of relative positional vector. Instead of adding the vectors to the key vectors in the self-attention layer, our method transforms the key vectors according to its position. Experiments on LibriSpeech dataset show that our approach outperforms the positional vector approach.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2021 12th International Symposium on Chinese Spoken Language Processing (ISCSLP)

自引率

0.00%

发文量