基于伪监督学习的连续手语识别

Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces Pub Date : 2019-10-15 DOI:10.1145/3347319.3356837

Xiankun Pei, Dan Guo, Ye Zhao

{"title":"基于伪监督学习的连续手语识别","authors":"Xiankun Pei, Dan Guo, Ye Zhao","doi":"10.1145/3347319.3356837","DOIUrl":null,"url":null,"abstract":"Continuous sign language recognition task is challenging for the reason that the ordered words have no exact temporal locations in the video. Aiming at this problem, we propose a method based on pseudo-supervised learning. First, we use a 3D residual convolutional network (3D-ResNet) pre-trained on the UCF101 dataset to extract visual features. Second, we employ a sequence model with connectionist temporal classification (CTC) loss for learning the mapping between the visual features and sentence-level labels, which can be used to generate clip-level pseudo-labels. Since the CTC objective function has limited effects on visual features extracted from early 3D-ResNet, we fine-tune the 3D-ResNet by feeding the clip-level pseudo-labels and video clips to obtain better feature representation. The feature extractor and the sequence model are optimized alternately with CTC loss. The effectiveness of the proposed method is verified on the large datasets RWTH-PHOENIX-Weather-2014.","PeriodicalId":420165,"journal":{"name":"Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces","volume":"145 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-10-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"Continuous Sign Language Recognition Based on Pseudo-supervised Learning\",\"authors\":\"Xiankun Pei, Dan Guo, Ye Zhao\",\"doi\":\"10.1145/3347319.3356837\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Continuous sign language recognition task is challenging for the reason that the ordered words have no exact temporal locations in the video. Aiming at this problem, we propose a method based on pseudo-supervised learning. First, we use a 3D residual convolutional network (3D-ResNet) pre-trained on the UCF101 dataset to extract visual features. Second, we employ a sequence model with connectionist temporal classification (CTC) loss for learning the mapping between the visual features and sentence-level labels, which can be used to generate clip-level pseudo-labels. Since the CTC objective function has limited effects on visual features extracted from early 3D-ResNet, we fine-tune the 3D-ResNet by feeding the clip-level pseudo-labels and video clips to obtain better feature representation. The feature extractor and the sequence model are optimized alternately with CTC loss. The effectiveness of the proposed method is verified on the large datasets RWTH-PHOENIX-Weather-2014.\",\"PeriodicalId\":420165,\"journal\":{\"name\":\"Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces\",\"volume\":\"145 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2019-10-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/3347319.3356837\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3347319.3356837","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

摘要

连续的手语识别任务具有挑战性，因为有序的单词在视频中没有确切的时间位置。针对这一问题，我们提出了一种基于伪监督学习的方法。首先，我们使用在UCF101数据集上预训练的3D残差卷积网络(3D- resnet)提取视觉特征。其次，我们采用具有连接时间分类(CTC)损失的序列模型来学习视觉特征与句子级标签之间的映射关系，该模型可用于生成剪辑级伪标签。由于CTC目标函数对早期3D-ResNet提取的视觉特征影响有限，我们通过输入片段级伪标签和视频片段对3D-ResNet进行微调，以获得更好的特征表示。利用CTC损失对特征提取器和序列模型交替优化。在RWTH-PHOENIX-Weather-2014大型数据集上验证了该方法的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Continuous Sign Language Recognition Based on Pseudo-supervised Learning

Continuous sign language recognition task is challenging for the reason that the ordered words have no exact temporal locations in the video. Aiming at this problem, we propose a method based on pseudo-supervised learning. First, we use a 3D residual convolutional network (3D-ResNet) pre-trained on the UCF101 dataset to extract visual features. Second, we employ a sequence model with connectionist temporal classification (CTC) loss for learning the mapping between the visual features and sentence-level labels, which can be used to generate clip-level pseudo-labels. Since the CTC objective function has limited effects on visual features extracted from early 3D-ResNet, we fine-tune the 3D-ResNet by feeding the clip-level pseudo-labels and video clips to obtain better feature representation. The feature extractor and the sequence model are optimized alternately with CTC loss. The effectiveness of the proposed method is verified on the large datasets RWTH-PHOENIX-Weather-2014.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings of the 2nd Workshop on Multimedia for Accessible Human Computer Interfaces

自引率

0.00%

发文量