Textual Datasets For Portuguese-Brazilian Language Models

Anais do IV Dataset Showcase Workshop (DSW 2022) Pub Date : 2022-09-19 DOI:10.5753/dsw.2022.224294

Matheus Ferraroni Sanches, Jáder M. C. de Sá, Henrique T. S. Foerste, R. R. Souza, J. C. dos Reis, L. Villas

引用次数: 1

Abstract

Advances in Natural Language Processing have generated new models that push forward the state of the art. This reached new heights in complex tasks in handling unstructured texts. Most of the new architectures and models focus on the English language. There is a lack of available datasets that can be used during the training of new models. This investigation presents four new textual datasets for language modeling in Brazilian Portuguese. Our datasets were generated from several specific methodologies that aimed to obtain data of different natures. Two of our sets were originally built from data in online web forums. We also distribute a translated version of MultiWOZ, and a clean version of BrWaC. The original datasets are made available in a structured way to facilitate their use during the training of NLP models, with questions, answers and conversations already identified.

查看原文本刊更多论文

葡萄牙-巴西语言模型的文本数据集

自然语言处理的进步产生了新的模型，推动了艺术的发展。这在处理非结构化文本的复杂任务中达到了新的高度。大多数新的架构和模型都集中在英语语言上。在训练新模型的过程中，缺乏可用的数据集。本研究提出了巴西葡萄牙语语言建模的四个新的文本数据集。我们的数据集是由几种特定的方法生成的，旨在获得不同性质的数据。我们的两个集合最初是根据在线网络论坛中的数据构建的。我们还发布了MultiWOZ的翻译版本和BrWaC的干净版本。原始数据集以结构化的方式提供，以便在NLP模型的训练过程中使用，其中已经确定了问题，答案和对话。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Anais do IV Dataset Showcase Workshop (DSW 2022)

自引率

0.00%

发文量