An Empirical Study for Vietnamese Constituency Parsing with Pre-training

2021 RIVF International Conference on Computing and Communication Technologies (RIVF) Pub Date : 2020-10-19 DOI:10.1109/RIVF51545.2021.9642143

Tuan-Vi Tran, Xuan-Thien Pham, Duc-Vu Nguyen, Kiet Van Nguyen, N. Nguyen

引用次数: 2

Abstract

Constituency parsing is an important task that gets more attention in natural language processing. In this work, we use a span-based approach for Vietnamese constituency parsing. Our method follows the self-attention encoder architecture and a chart decoder using a CKY-style inference algorithm. We present analyses of the experiment results of the comparison of our empirical method using pre-training models XLM-R and PhoBERT on both Vietnamese datasets VietTreebank and NIIVTB1. The results show that our model with XLM-R archived the significantly F1-score better than other pre-training models, VietTreebank at 81.19% and NIIVTB1 at 85.70%.

查看原文本刊更多论文

基于预训练的越南语选区分析实证研究

成分分析是自然语言处理中备受关注的一项重要任务。在这项工作中，我们使用基于跨度的方法进行越南语选区解析。我们的方法遵循自关注编码器架构和使用cky风格推理算法的图表解码器。我们在越南数据集VietTreebank和NIIVTB1上使用预训练模型XLM-R和PhoBERT对我们的经验方法的实验结果进行了比较分析。结果表明，我们的XLM-R模型的f1得分明显优于其他预训练模型，VietTreebank为81.19%，NIIVTB1为85.70%。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2021 RIVF International Conference on Computing and Communication Technologies (RIVF)

自引率

0.00%

发文量