Continuous Sign Language Interpretation to Text Using Deep Learning Models

2022 25th International Conference on Computer and Information Technology (ICCIT) Pub Date : 2022-12-17 DOI:10.1109/ICCIT57492.2022.10054721

Afridi Ibn Rahman, Zebel-E.-Noor Akhand, Tasin Al Nahian Khan, Anirudh Sarda, Subhi Bhuiyan, Mma Rakib, Zubayer Ahmed Fahim, Indronil Kundu

{"title":"Continuous Sign Language Interpretation to Text Using Deep Learning Models","authors":"Afridi Ibn Rahman, Zebel-E.-Noor Akhand, Tasin Al Nahian Khan, Anirudh Sarda, Subhi Bhuiyan, Mma Rakib, Zubayer Ahmed Fahim, Indronil Kundu","doi":"10.1109/ICCIT57492.2022.10054721","DOIUrl":null,"url":null,"abstract":"The COVID-19 pandemic has obligated people to adopt the virtual lifestyle. Currently, the use of videoconferencing to conduct business meetings is prevalent owing to the numerous benefits it presents. However, a large number of people with speech impediment find themselves handicapped to the new normal as they cannot communicate their ideas effectively, especially in fast paced meetings. Therefore, this paper aims to introduce an enriched dataset using an action recognition method with the most common phrases translated into American Sign Language (ASL) that are routinely used in professional meetings. It further proposes a sign language detecting and classifying model employing deep learning architectures, namely, CNN and LSTM. The performances of these models are analysed by employing different performance metrics like accuracy, recall, F1- Score and Precision. CNN and LSTM models yield an accuracy of 93.75% and 96.54% respectively, after being trained with the dataset introduced in this study. Therefore, the incorporation of the LSTM model into different cloud services, virtual private networks and softwares will allow people with speech impairment to use sign language, which will automatically be translated into captions using moving camera circumstances in real time. This will in turn equip other people with the tool to understand and grasp the message that is being conveyed and easily discuss and effectuate the ideas.","PeriodicalId":255498,"journal":{"name":"2022 25th International Conference on Computer and Information Technology (ICCIT)","volume":"52 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-12-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 25th International Conference on Computer and Information Technology (ICCIT)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICCIT57492.2022.10054721","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

The COVID-19 pandemic has obligated people to adopt the virtual lifestyle. Currently, the use of videoconferencing to conduct business meetings is prevalent owing to the numerous benefits it presents. However, a large number of people with speech impediment find themselves handicapped to the new normal as they cannot communicate their ideas effectively, especially in fast paced meetings. Therefore, this paper aims to introduce an enriched dataset using an action recognition method with the most common phrases translated into American Sign Language (ASL) that are routinely used in professional meetings. It further proposes a sign language detecting and classifying model employing deep learning architectures, namely, CNN and LSTM. The performances of these models are analysed by employing different performance metrics like accuracy, recall, F1- Score and Precision. CNN and LSTM models yield an accuracy of 93.75% and 96.54% respectively, after being trained with the dataset introduced in this study. Therefore, the incorporation of the LSTM model into different cloud services, virtual private networks and softwares will allow people with speech impairment to use sign language, which will automatically be translated into captions using moving camera circumstances in real time. This will in turn equip other people with the tool to understand and grasp the message that is being conveyed and easily discuss and effectuate the ideas.

查看原文本刊更多论文

使用深度学习模型的连续手语文本解释

新冠肺炎疫情迫使人们采用虚拟生活方式。目前，使用视频会议进行商务会议是普遍的，因为它提供了许多好处。然而，大量有语言障碍的人发现自己无法适应新常态，因为他们无法有效地表达自己的想法，尤其是在快节奏的会议中。因此，本文旨在引入一个丰富的数据集，使用一种动作识别方法，将最常见的短语翻译成专业会议中经常使用的美国手语(ASL)。在此基础上，提出了一种基于CNN和LSTM深度学习的手语检测与分类模型。通过采用不同的性能指标，如准确率、召回率、F1- Score和精度，分析了这些模型的性能。CNN和LSTM模型使用本文引入的数据集进行训练后，准确率分别为93.75%和96.54%。因此，将LSTM模型整合到不同的云服务、虚拟专用网络和软件中，将允许有语言障碍的人使用手语，这些手语将通过移动的摄像机环境实时自动翻译成字幕。这将反过来为其他人提供理解和掌握所传达的信息的工具，并容易讨论和实现这些想法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2022 25th International Conference on Computer and Information Technology (ICCIT)

自引率

0.00%

发文量