Speech recognition in adverse conditions by humans and machines.

IF 1.2 Q3 ACOUSTICS

JASA express letters Pub Date : 2024-11-01 DOI:10.1121/10.0032473

Chloe Patman, Eleanor Chodroff

引用次数: 0

Abstract

In the development of automatic speech recognition systems, achieving human-like performance has been a long-held goal. Recent releases of large spoken language models have claimed to achieve such performance, although direct comparison to humans has been severely limited. The present study tested L1 British English listeners against two automatic speech recognition systems (wav2vec 2.0 and Whisper, base and large sizes) in adverse listening conditions: speech-shaped noise and pub noise, at different signal-to-noise ratios, and recordings produced with or without face masks. Humans maintained the advantage against all systems, except for Whisper large, which outperformed humans in every condition but pub noise.

查看原文本刊更多论文

人类和机器在恶劣条件下的语音识别。

在自动语音识别系统的开发过程中，实现与人类相似的性能是一个长期坚持的目标。最近发布的大型口语模型声称可以达到这样的性能，但与人类的直接比较却受到严重限制。本研究针对两个自动语音识别系统（wav2vec 2.0 和 Whisper，基本型和大型），在不利的听力条件下对英国英语中级听者进行了测试：不同信噪比下的语音噪声和酒吧噪声，以及带或不带面罩的录音。人类在所有系统中都保持优势，只有 Whisper large 除酒吧噪音外在所有条件下都优于人类。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

JASA express letters

CiteScore

1.70

自引率

0.00%

发文量