Exploring Continual Learning for Code Generation Models

Annual Meeting of the Association for Computational Linguistics Pub Date : 2023-07-05 DOI:10.48550/arXiv.2307.02435

Prateek Yadav, Q. Sun, Hantian Ding, Xiaopeng Li, Dejiao Zhang, Ming Tan, Xiaofei Ma, Parminder Bhatia, Ramesh Nallapati, M. Ramanathan, Mohit Bansal, Bing Xiang

{"title":"Exploring Continual Learning for Code Generation Models","authors":"Prateek Yadav, Q. Sun, Hantian Ding, Xiaopeng Li, Dejiao Zhang, Ming Tan, Xiaofei Ma, Parminder Bhatia, Ramesh Nallapati, M. Ramanathan, Mohit Bansal, Bing Xiang","doi":"10.48550/arXiv.2307.02435","DOIUrl":null,"url":null,"abstract":"Large-scale code generation models such as Copilot and CodeT5 have achieved impressive performance. However, libraries are upgraded or deprecated very frequently and re-training large-scale language models is computationally expensive. Therefore, Continual Learning (CL) is an important aspect that remains under-explored in the code domain. In this paper, we introduce a benchmark called CodeTask-CL that covers a wide range of tasks, including code generation, translation, summarization, and refinement, with different input and output programming languages. Next, on our CodeTask-CL benchmark, we compare popular CL techniques from NLP and Vision domains. We find that effective methods like Prompt Pooling (PP) suffer from catastrophic forgetting due to the unstable training of the prompt selection mechanism caused by stark distribution shifts in coding tasks. We address this issue with our proposed method, Prompt Pooling with Teacher Forcing (PP-TF), that stabilizes training by enforcing constraints on the prompt selection mechanism and leads to a 21.54% improvement over Prompt Pooling. Along with the benchmark, we establish a training pipeline that can be used for CL on code models, which we believe can motivate further development of CL methods for code models.","PeriodicalId":352845,"journal":{"name":"Annual Meeting of the Association for Computational Linguistics","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-07-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Annual Meeting of the Association for Computational Linguistics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.48550/arXiv.2307.02435","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

Abstract

Large-scale code generation models such as Copilot and CodeT5 have achieved impressive performance. However, libraries are upgraded or deprecated very frequently and re-training large-scale language models is computationally expensive. Therefore, Continual Learning (CL) is an important aspect that remains under-explored in the code domain. In this paper, we introduce a benchmark called CodeTask-CL that covers a wide range of tasks, including code generation, translation, summarization, and refinement, with different input and output programming languages. Next, on our CodeTask-CL benchmark, we compare popular CL techniques from NLP and Vision domains. We find that effective methods like Prompt Pooling (PP) suffer from catastrophic forgetting due to the unstable training of the prompt selection mechanism caused by stark distribution shifts in coding tasks. We address this issue with our proposed method, Prompt Pooling with Teacher Forcing (PP-TF), that stabilizes training by enforcing constraints on the prompt selection mechanism and leads to a 21.54% improvement over Prompt Pooling. Along with the benchmark, we establish a training pipeline that can be used for CL on code models, which we believe can motivate further development of CL methods for code models.

查看原文本刊更多论文

探索代码生成模型的持续学习

像Copilot和CodeT5这样的大规模代码生成模型已经取得了令人印象深刻的性能。然而，库的升级或弃用非常频繁，并且重新训练大规模语言模型在计算上是昂贵的。因此，持续学习(CL)是代码领域中尚未得到充分开发的一个重要方面。在本文中，我们介绍了一个名为CodeTask-CL的基准测试，它涵盖了使用不同输入和输出编程语言的各种任务，包括代码生成、翻译、汇总和细化。接下来，在我们的codettask -CL基准测试中，我们比较了来自NLP和视觉领域的流行CL技术。我们发现，由于编码任务的明显分布变化导致提示选择机制训练不稳定，提示池(Prompt Pooling, PP)等有效方法容易出现灾难性遗忘。我们提出了一种方法来解决这个问题，即基于教师强迫的提示池(PP-TF)，该方法通过对提示选择机制施加约束来稳定培训，比提示池提高了21.54%。除了这个基准，我们还建立了一个训练管道，可以用于代码模型上的CL，我们相信这可以激励代码模型上的CL方法的进一步开发。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Annual Meeting of the Association for Computational Linguistics

自引率

0.00%

发文量