Data Extraction from Similarly Structured Scanned Documents

Russian Digital Libraries Journal Pub Date : 2021-09-12 DOI:10.26907/1562-5419-2021-24-4-667-688

Rustem Damirovich Saitgareev, B. R. Giniatullin, Vladislav Yurievich Toporov, Artur Aleksandrovich Atnagulov, Farid Radikovich Aglyamov

引用次数: 0

Abstract

Currently, the major part of transmitted and stored data is unstructured, and the amount of unstructured data is growing rapidly each year, although it is hardly searchable, unqueryable, and its processing is not automated. At the same time, there is a growth of electronic document management systems. This paper proposes a solution for extracting data from paper documents considering their structure and layout based on document photos. By examining different approaches, including neural networks and plain algorithmic methods, we present their results and discuss them.

查看原文本刊更多论文

从结构相似的扫描文档中提取数据

目前，传输和存储的大部分数据是非结构化的，非结构化数据的数量每年都在快速增长，尽管它很难被搜索、查询，而且它的处理也不是自动化的。与此同时，电子文档管理系统也在不断发展。本文提出了一种基于文档照片的基于文档结构和布局的纸质文档数据提取方法。通过研究不同的方法，包括神经网络和普通算法方法，我们提出了他们的结果并讨论了他们。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Russian Digital Libraries Journal

自引率

0.00%

发文量