在Zooniverse平台上结合人类和机器转录

NUT@EMNLP Pub Date : 2018-11-01 DOI:10.18653/v1/W18-6129

Daniel Hanson, A. Simenstad

{"title":"在Zooniverse平台上结合人类和机器转录","authors":"Daniel Hanson, A. Simenstad","doi":"10.18653/v1/W18-6129","DOIUrl":null,"url":null,"abstract":"Transcribing handwritten documents to create fully searchable texts is an essential part of the archival process. Traditional text recognition methods, such as optical character recognition (OCR), do not work on handwritten documents due to their frequent noisiness and OCR’s need for individually segmented letters. Crowdsourcing and improved machine models are two modern methods for transcribing handwritten documents.","PeriodicalId":207795,"journal":{"name":"NUT@EMNLP","volume":"5 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2018-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"Combining Human and Machine Transcriptions on the Zooniverse Platform\",\"authors\":\"Daniel Hanson, A. Simenstad\",\"doi\":\"10.18653/v1/W18-6129\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Transcribing handwritten documents to create fully searchable texts is an essential part of the archival process. Traditional text recognition methods, such as optical character recognition (OCR), do not work on handwritten documents due to their frequent noisiness and OCR’s need for individually segmented letters. Crowdsourcing and improved machine models are two modern methods for transcribing handwritten documents.\",\"PeriodicalId\":207795,\"journal\":{\"name\":\"NUT@EMNLP\",\"volume\":\"5 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2018-11-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"NUT@EMNLP\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.18653/v1/W18-6129\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"NUT@EMNLP","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/W18-6129","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

转录手写文件以创建完全可搜索的文本是存档过程的重要组成部分。传统的文本识别方法，如光学字符识别(OCR)，由于手写文档的频繁噪声和OCR需要单独分割字母，不能用于手写文档。众包和改进的机器模型是抄写手写文件的两种现代方法。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Combining Human and Machine Transcriptions on the Zooniverse Platform

Transcribing handwritten documents to create fully searchable texts is an essential part of the archival process. Traditional text recognition methods, such as optical character recognition (OCR), do not work on handwritten documents due to their frequent noisiness and OCR’s need for individually segmented letters. Crowdsourcing and improved machine models are two modern methods for transcribing handwritten documents.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

NUT@EMNLP

自引率

0.00%

发文量