Continual Learning for Natural Language Generations with Transformer Calibration

Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL) Pub Date : 1900-01-01 DOI:10.18653/v1/2022.conll-1.4

Peng Yang, Dingcheng Li, Ping Li

{"title":"Continual Learning for Natural Language Generations with Transformer Calibration","authors":"Peng Yang, Dingcheng Li, Ping Li","doi":"10.18653/v1/2022.conll-1.4","DOIUrl":null,"url":null,"abstract":"Conventional natural language process (NLP) generation models are trained offline with a given dataset for a particular task, which is referred to as isolated learning. Research on sequence-to-sequence language generation aims to study continual learning model to constantly learning from sequentially encountered tasks. However, continual learning studies often suffer from catastrophic forgetting, a persistent challenge for lifelong learning. In this paper, we present a novel NLP transformer model that attempts to mitigate catastrophic forgetting in online continual learning from a new perspective, i.e., attention calibration. We model the attention in the transformer as a calibrated unit in a general formulation, where the attention calibration could give benefits to balance the stability and plasticity of continual learning algorithms through influencing both their forward inference path and backward optimization path. Our empirical experiments, paraphrase generation and dialog response generation, demonstrate that this work outperforms state-of-the-art models by a considerable margin and effectively mitigate the forgetting.","PeriodicalId":221345,"journal":{"name":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","volume":"59 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"3","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/2022.conll-1.4","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 3

Abstract

Conventional natural language process (NLP) generation models are trained offline with a given dataset for a particular task, which is referred to as isolated learning. Research on sequence-to-sequence language generation aims to study continual learning model to constantly learning from sequentially encountered tasks. However, continual learning studies often suffer from catastrophic forgetting, a persistent challenge for lifelong learning. In this paper, we present a novel NLP transformer model that attempts to mitigate catastrophic forgetting in online continual learning from a new perspective, i.e., attention calibration. We model the attention in the transformer as a calibrated unit in a general formulation, where the attention calibration could give benefits to balance the stability and plasticity of continual learning algorithms through influencing both their forward inference path and backward optimization path. Our empirical experiments, paraphrase generation and dialog response generation, demonstrate that this work outperforms state-of-the-art models by a considerable margin and effectively mitigate the forgetting.

查看原文本刊更多论文

变压器校准自然语言世代的持续学习

传统的自然语言过程(NLP)生成模型是针对特定任务使用给定数据集离线训练的，这被称为孤立学习。序列到序列语言生成研究的目的是研究连续学习模型，从顺序遇到的任务中不断学习。然而，持续的学习往往会遭受灾难性的遗忘，这是终身学习的一个持续挑战。在本文中，我们提出了一种新的NLP变压器模型，该模型试图从一个新的角度，即注意校准，来减轻在线持续学习中的灾难性遗忘。我们将变压器中的注意力建模为一般公式中的校准单元，其中注意力校准可以通过影响其前向推理路径和后向优化路径来平衡持续学习算法的稳定性和可塑性。我们的经验实验，释义生成和对话响应生成，证明了这项工作在相当程度上优于最先进的模型，并有效地减轻了遗忘。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL)

自引率

0.00%

发文量