打印文档图像的压缩和字符串匹配方法

2009 10th International Conference on Document Analysis and Recognition Pub Date : 2009-07-26 DOI:10.1109/ICDAR.2009.182

Hajime Imura, Yuzuru Tanaka

{"title":"打印文档图像的压缩和字符串匹配方法","authors":"Hajime Imura, Yuzuru Tanaka","doi":"10.1109/ICDAR.2009.182","DOIUrl":null,"url":null,"abstract":"This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \\& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.","PeriodicalId":433762,"journal":{"name":"2009 10th International Conference on Document Analysis and Recognition","volume":"66 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2009-07-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"Compression and String Matching Method for Printed Document Images\",\"authors\":\"Hajime Imura, Yuzuru Tanaka\",\"doi\":\"10.1109/ICDAR.2009.182\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \\\\& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.\",\"PeriodicalId\":433762,\"journal\":{\"name\":\"2009 10th International Conference on Document Analysis and Recognition\",\"volume\":\"66 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2009-07-26\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2009 10th International Conference on Document Analysis and Recognition\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICDAR.2009.182\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2009 10th International Conference on Document Analysis and Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDAR.2009.182","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

摘要

本文介绍了一种打印文档图像的压缩技术和压缩图像的字符串匹配方法。为了通过Web发送数字化文档图像，需要对文档图像进行压缩。此外，为了处理历史凸版印刷馆藏，提供全文检索方法是很重要的。提出的压缩方案是基于字符模式匹配&替换方法，使用文档图像的字符串匹配技术。该方法采用基于统计字符形状特征的伪编码，不受语言和字体差异的影响。我们还在压缩文档的字符串匹配中使用伪代码。该系统与机器可读文本的全文搜索一样快。我们的方法在压缩大小下进行了评估，计算了基于n-gram的查询字符串的召回精度曲线。实验表明，大约100页的300 dpi灰度文档可以压缩到1兆字节左右。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Compression and String Matching Method for Printed Document Images

This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2009 10th International Conference on Document Analysis and Recognition

自引率

0.00%

发文量