A multimodal approach to initialisation for top-down speaker diarization of television shows

2010 18th European Signal Processing Conference Pub Date : 2010-08-23 DOI:10.5281/ZENODO.41986

Simon Bozonnet, Félicien Vallet, N. Evans, S. Essid, G. Richard, J. Carrive

引用次数: 12

Abstract

This paper presents a new multimodal approach to speaker diarization of TV show data. We hypothesize that the intraspeaker variation in visual information might be less than that in the corresponding acoustic information and therefore might be better suited to the task of speaker model initialisation. This is an acknowledged weakness of the computationally efficient top-down approach to speaker diarization that is used here. Experimental results show that a recently proposed approach to purification and the new multimodal approach to initialisation together deliver 22% and 17% relative improvements in diarization performance over the baseline system on independent development and evaluation datasets respectively.

查看原文本刊更多论文

电视节目自顶向下说话人初始化的多模态方法

本文提出了一种新的多模态电视节目数据说话人特征化方法。我们假设说话人内部视觉信息的变化可能小于相应的声学信息的变化，因此可能更适合说话人模型初始化的任务。这是这里使用的计算效率高的自顶向下的说话人特征化方法的公认弱点。实验结果表明，最近提出的净化方法和新的多模态初始化方法在独立开发和评估数据集上分别比基线系统的初始化性能提高22%和17%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2010 18th European Signal Processing Conference

自引率

0.00%

发文量