{"title":"打印文档图像的压缩和字符串匹配方法","authors":"Hajime Imura, Yuzuru Tanaka","doi":"10.1109/ICDAR.2009.182","DOIUrl":null,"url":null,"abstract":"This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \\& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.","PeriodicalId":433762,"journal":{"name":"2009 10th International Conference on Document Analysis and Recognition","volume":"66 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2009-07-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"Compression and String Matching Method for Printed Document Images\",\"authors\":\"Hajime Imura, Yuzuru Tanaka\",\"doi\":\"10.1109/ICDAR.2009.182\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \\\\& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.\",\"PeriodicalId\":433762,\"journal\":{\"name\":\"2009 10th International Conference on Document Analysis and Recognition\",\"volume\":\"66 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2009-07-26\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2009 10th International Conference on Document Analysis and Recognition\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICDAR.2009.182\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2009 10th International Conference on Document Analysis and Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDAR.2009.182","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Compression and String Matching Method for Printed Document Images
This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character Pattern Matching \& Substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.