Audiovisual-based adaptive speaker identification

2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03). Pub Date : 2003-07-06 DOI:10.1109/ICASSP.2003.1200095

Y. Li, Shrikanth S. Narayanan, C.-C. Jay Kuo

引用次数: 1

Abstract

An adaptive speaker identification system is presented in this paper, which aims to recognize speakers in feature films by exploiting both audio and visual cues. Specifically, the audio source is first analyzed to identify speakers using a likelihood-based approach. Meanwhile, the visual source is parsed to recognize talking faces using face detection/recognition and mouth tracking techniques. These two information sources are then integrated under a probabilistic framework for improved system performance. Moreover, to account for speakers' voice variations along time, we update their acoustic models on the fly by adapting to their newly contributed speech data. An average of 80% identification accuracy has been achieved on two test movies. This shows a promising future for the proposed audiovisual-based adaptive speaker identification approach.

查看原文本刊更多论文

基于视听的自适应说话人识别

本文提出了一种自适应说话人识别系统，该系统旨在利用视觉和听觉线索对故事片中的说话人进行识别。具体来说，首先分析音频源以使用基于可能性的方法识别说话者。同时，利用人脸检测/识别和嘴部跟踪技术对视觉源进行解析，实现说话人脸的识别。然后将这两个信息源集成到一个概率框架下，以改进系统性能。此外，为了解释说话者的声音随时间的变化，我们通过适应他们新提供的语音数据来更新他们的声学模型。在两个测试影片上平均达到了80%的识别准确率。这表明基于视听的自适应说话人识别方法具有广阔的应用前景。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).

自引率

0.00%

发文量