一个基准数据集的叙事学生散文与多能力等级自动作文评分在巴西葡萄牙语

IF 1 Q3 MULTIDISCIPLINARY SCIENCES
Hilário Oliveira , Rafael Ferreira Mello , Péricles Miranda , Hyan Batista , Moésio Wenceslau da Silva Filho , Thiago Cordeiro , Ig Ibert Bittencourt , Seiji Isotani
{"title":"一个基准数据集的叙事学生散文与多能力等级自动作文评分在巴西葡萄牙语","authors":"Hilário Oliveira ,&nbsp;Rafael Ferreira Mello ,&nbsp;Péricles Miranda ,&nbsp;Hyan Batista ,&nbsp;Moésio Wenceslau da Silva Filho ,&nbsp;Thiago Cordeiro ,&nbsp;Ig Ibert Bittencourt ,&nbsp;Seiji Isotani","doi":"10.1016/j.dib.2025.111526","DOIUrl":null,"url":null,"abstract":"<div><div>This paper describes the development of a new database comprising 1235 narrative essays written in Portuguese by 5th-grade students in Brazil. The corpus construction process involved three main steps: acquiring and transcribing photos of the essays, annotating them based on a real pre-defined correction rubric by experts considering four key writing competencies (formal language use, textual typology, thematic coherence, and textual cohesion), and resolving disagreements between the annotators. Two human experts manually evaluated each essay using a five-point scale (Level I: Complete lack of domain - Level V: Excellent mastery) aligned with the correction rubric. In cases of disagreement between the initial evaluators, a third expert facilitated the divergences resolution. To the best of our knowledge, this is the first publicly available dataset of elementary school essays in Brazilian Portuguese that features narrative writing samples with corresponding grades across multiple competencies commonly used in writing assessment. We believe this resource can contribute to developing automatic essay scoring systems tailored for evaluating narrative texts written in Brazilian Portuguese.</div></div>","PeriodicalId":10973,"journal":{"name":"Data in Brief","volume":"60 ","pages":"Article 111526"},"PeriodicalIF":1.0000,"publicationDate":"2025-03-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"A benchmark dataset of narrative student essays with multi-competency grades for automatic essay scoring in Brazilian Portuguese\",\"authors\":\"Hilário Oliveira ,&nbsp;Rafael Ferreira Mello ,&nbsp;Péricles Miranda ,&nbsp;Hyan Batista ,&nbsp;Moésio Wenceslau da Silva Filho ,&nbsp;Thiago Cordeiro ,&nbsp;Ig Ibert Bittencourt ,&nbsp;Seiji Isotani\",\"doi\":\"10.1016/j.dib.2025.111526\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><div>This paper describes the development of a new database comprising 1235 narrative essays written in Portuguese by 5th-grade students in Brazil. The corpus construction process involved three main steps: acquiring and transcribing photos of the essays, annotating them based on a real pre-defined correction rubric by experts considering four key writing competencies (formal language use, textual typology, thematic coherence, and textual cohesion), and resolving disagreements between the annotators. Two human experts manually evaluated each essay using a five-point scale (Level I: Complete lack of domain - Level V: Excellent mastery) aligned with the correction rubric. In cases of disagreement between the initial evaluators, a third expert facilitated the divergences resolution. To the best of our knowledge, this is the first publicly available dataset of elementary school essays in Brazilian Portuguese that features narrative writing samples with corresponding grades across multiple competencies commonly used in writing assessment. We believe this resource can contribute to developing automatic essay scoring systems tailored for evaluating narrative texts written in Brazilian Portuguese.</div></div>\",\"PeriodicalId\":10973,\"journal\":{\"name\":\"Data in Brief\",\"volume\":\"60 \",\"pages\":\"Article 111526\"},\"PeriodicalIF\":1.0000,\"publicationDate\":\"2025-03-27\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Data in Brief\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S2352340925002586\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q3\",\"JCRName\":\"MULTIDISCIPLINARY SCIENCES\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Data in Brief","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2352340925002586","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"MULTIDISCIPLINARY SCIENCES","Score":null,"Total":0}
引用次数: 0

摘要

本文描述了巴西五年级学生用葡萄牙语写的1235篇叙事文章的新数据库的开发。语料库构建过程包括三个主要步骤:获取和转录文章的照片,由专家根据四个关键写作能力(正式语言使用、文本类型、主题连贯和文本衔接)根据真正预先定义的纠正规则对其进行注释,并解决注释者之间的分歧。两名人类专家使用五分制(一级:完全缺乏领域- V级:优秀掌握)与纠正规则对齐,手动评估每篇文章。如果最初评价人员之间存在分歧,第三位专家将协助解决分歧。据我们所知,这是第一个公开可用的巴西葡萄牙语小学论文数据集,其特点是叙事写作样本具有相应的等级,涵盖写作评估中常用的多种能力。我们相信这个资源可以有助于开发自动论文评分系统量身定制的评估叙事文本写在巴西葡萄牙语。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
A benchmark dataset of narrative student essays with multi-competency grades for automatic essay scoring in Brazilian Portuguese
This paper describes the development of a new database comprising 1235 narrative essays written in Portuguese by 5th-grade students in Brazil. The corpus construction process involved three main steps: acquiring and transcribing photos of the essays, annotating them based on a real pre-defined correction rubric by experts considering four key writing competencies (formal language use, textual typology, thematic coherence, and textual cohesion), and resolving disagreements between the annotators. Two human experts manually evaluated each essay using a five-point scale (Level I: Complete lack of domain - Level V: Excellent mastery) aligned with the correction rubric. In cases of disagreement between the initial evaluators, a third expert facilitated the divergences resolution. To the best of our knowledge, this is the first publicly available dataset of elementary school essays in Brazilian Portuguese that features narrative writing samples with corresponding grades across multiple competencies commonly used in writing assessment. We believe this resource can contribute to developing automatic essay scoring systems tailored for evaluating narrative texts written in Brazilian Portuguese.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
Data in Brief
Data in Brief MULTIDISCIPLINARY SCIENCES-
CiteScore
3.10
自引率
0.00%
发文量
996
审稿时长
70 days
期刊介绍: Data in Brief provides a way for researchers to easily share and reuse each other''s datasets by publishing data articles that: -Thoroughly describe your data, facilitating reproducibility. -Make your data, which is often buried in supplementary material, easier to find. -Increase traffic towards associated research articles and data, leading to more citations. -Open up doors for new collaborations. Because you never know what data will be useful to someone else, Data in Brief welcomes submissions that describe data from all research areas.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信