From Text Detection in Videos to Person Identification

2012 IEEE International Conference on Multimedia and Expo Pub Date : 2012-07-09 DOI:10.1109/ICME.2012.119

Johann Poignant, L. Besacier, G. Quénot, F. Thollard

引用次数: 53

Abstract

We present in this article a video OCR system that detects and recognizes overlaid texts in video as well as its application to person identification in video documents. We proceed in several steps. First, text detection and temporal tracking are performed. After adaptation of images to a standard OCR system, a final post-processing combines multiple transcriptions of the same text box. The semi-supervised adaptation of this system to a particular video type (video broadcast from a French TV) is proposed and evaluated. The system is efficient as it runs 3 times faster than real time (including the OCR step) on a desktop Linux box. Both text detection and recognition are evaluated individually and through a person recognition task where it is shown that the combination of OCR and audio (speaker) information can greatly improve the performances of a state of the art audio based person identification system.

查看原文本刊更多论文

从视频文本检测到人物识别

本文介绍了一种视频OCR系统，用于检测和识别视频中的覆盖文本，并将其应用于视频文档中的人物识别。我们分几个步骤进行。首先，进行文本检测和时间跟踪。在将图像适应标准OCR系统之后，最后的后处理将同一文本框的多个转录组合在一起。提出并评估了该系统对特定视频类型(来自法国电视的视频广播)的半监督适应。该系统非常高效，因为它在桌面Linux上的运行速度比实时(包括OCR步骤)快3倍。文本检测和识别分别通过一个人识别任务进行评估，其中显示OCR和音频(说话人)信息的组合可以大大提高基于音频的最先进的人识别系统的性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2012 IEEE International Conference on Multimedia and Expo

自引率

0.00%

发文量