Middle Zone Component Extraction and Recognition of Telugu Document Image

Ninth International Conference on Document Analysis and Recognition (ICDAR 2007) Pub Date : 2007-09-23 DOI:10.1109/ICDAR.2007.169

L. Reddy, L. Satyaprasad, A. Sastry

引用次数: 11

Abstract

Telugu is one of the ancient languages of South India. It has a complex orthography with a large number of distinct character shapes composed of simple and compound characters. The work reported in literature till the recent period is based on the connected component approach. Less attention is observed on the generalized character model and its application in the OCR development. Script syllable follows canonical structure where a consonant vowel core is preceded by one or two optional consonants .Formation of a syllable posses unique structural nature. In the present work, structural features of the syllable and the component model are combined to extract middle zone components. The shape of the middle zone components is closely related to a circle whereas other components are found with different topological features. Recognition rate of 99 percent is observed with the proposed method.

查看原文本刊更多论文

泰卢固语文档图像中间区分量提取与识别

泰卢固语是印度南部的一种古老语言。它有一个复杂的正字法，有大量不同的汉字形状，由单字和复合字组成。直到最近，文献报道的工作都是基于关联成分方法。一般对广义字符模型及其在OCR开发中的应用关注较少。手写体音节遵循规范结构，其中辅音元音核心前面有一个或两个可选的辅音。音节的形成具有独特的结构性质。本文采用音节结构特征和成分模型相结合的方法提取中间区成分。中间区成分的形状与圆密切相关，而其他成分具有不同的拓扑特征。该方法的识别率达到99%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Ninth International Conference on Document Analysis and Recognition (ICDAR 2007)

自引率

0.00%

发文量