后ocr文本处理的神经机器翻译方法

2022 30th Signal Processing and Communications Applications Conference (SIU) Pub Date : 2022-05-15 DOI:10.1109/SIU55565.2022.9864878

Ayse Irem Topcu, B. U. Töreyin

{"title":"后ocr文本处理的神经机器翻译方法","authors":"Ayse Irem Topcu, B. U. Töreyin","doi":"10.1109/SIU55565.2022.9864878","DOIUrl":null,"url":null,"abstract":"Optical Character Recognition (OCR) is the process of extracting the texts from the images by means of some special programs and transferring them to the computer environment. OCR quality directly affects the quality of most natural language processing processes. Many applications such as text classification, information extraction, text summarization with texts extracted from images are used in daily life. Therefore, detecting and correcting incorrectly translated texts after OCR is a topic that researchers are working on with many methods today. In this study, it is aimed to apply and observe the results on the dataset presented in the International Conference on Document Analysis and Recognition (ICDAR) 2019 OCR Post Error Detection and Correction competition, using the latest neural machine translation methods to find and correct post-OCR text errors.","PeriodicalId":115446,"journal":{"name":"2022 30th Signal Processing and Communications Applications Conference (SIU)","volume":"37 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-05-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Neural Machine Translation Approaches for Post-OCR Text Processing\",\"authors\":\"Ayse Irem Topcu, B. U. Töreyin\",\"doi\":\"10.1109/SIU55565.2022.9864878\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Optical Character Recognition (OCR) is the process of extracting the texts from the images by means of some special programs and transferring them to the computer environment. OCR quality directly affects the quality of most natural language processing processes. Many applications such as text classification, information extraction, text summarization with texts extracted from images are used in daily life. Therefore, detecting and correcting incorrectly translated texts after OCR is a topic that researchers are working on with many methods today. In this study, it is aimed to apply and observe the results on the dataset presented in the International Conference on Document Analysis and Recognition (ICDAR) 2019 OCR Post Error Detection and Correction competition, using the latest neural machine translation methods to find and correct post-OCR text errors.\",\"PeriodicalId\":115446,\"journal\":{\"name\":\"2022 30th Signal Processing and Communications Applications Conference (SIU)\",\"volume\":\"37 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-05-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2022 30th Signal Processing and Communications Applications Conference (SIU)\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/SIU55565.2022.9864878\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 30th Signal Processing and Communications Applications Conference (SIU)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SIU55565.2022.9864878","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

光学字符识别(OCR)是通过一些特殊的程序从图像中提取文本并将其传输到计算机环境中的过程。OCR的质量直接影响到大多数自然语言处理过程的质量。文本分类、信息提取、从图像中提取文本的文本摘要等应用在日常生活中得到了广泛的应用。因此，在OCR后对翻译错误的文本进行检测和纠正是目前研究人员正在用多种方法研究的课题。在本研究中，旨在应用和观察在国际文档分析与识别会议(ICDAR) 2019 OCR后错误检测和纠正竞赛中发表的数据集上的结果，使用最新的神经机器翻译方法来发现和纠正OCR后文本错误。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Neural Machine Translation Approaches for Post-OCR Text Processing

Optical Character Recognition (OCR) is the process of extracting the texts from the images by means of some special programs and transferring them to the computer environment. OCR quality directly affects the quality of most natural language processing processes. Many applications such as text classification, information extraction, text summarization with texts extracted from images are used in daily life. Therefore, detecting and correcting incorrectly translated texts after OCR is a topic that researchers are working on with many methods today. In this study, it is aimed to apply and observe the results on the dataset presented in the International Conference on Document Analysis and Recognition (ICDAR) 2019 OCR Post Error Detection and Correction competition, using the latest neural machine translation methods to find and correct post-OCR text errors.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2022 30th Signal Processing and Communications Applications Conference (SIU)

自引率

0.00%

发文量