如何比较没有目标长度的摘要?神经摘要文献的陷阱、解决方法与再审视

Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation Pub Date : 1900-01-01 DOI:10.18653/v1/W19-2303

Simeng Sun, Ori Shapira, Ido Dagan, A. Nenkova

{"title":"如何比较没有目标长度的摘要?神经摘要文献的陷阱、解决方法与再审视","authors":"Simeng Sun, Ori Shapira, Ido Dagan, A. Nenkova","doi":"10.18653/v1/W19-2303","DOIUrl":null,"url":null,"abstract":"We show that plain ROUGE F1 scores are not ideal for comparing current neural systems which on average produce different lengths. This is due to a non-linear pattern between ROUGE F1 and summary length. To alleviate the effect of length during evaluation, we have proposed a new method which normalizes the ROUGE F1 scores of a system by that of a random system with same average output length. A pilot human evaluation has shown that humans prefer short summaries in terms of the verbosity of a summary but overall consider longer summaries to be of higher quality. While human evaluations are more expensive in time and resources, it is clear that normalization, such as the one we proposed for automatic evaluation, will make human evaluations more meaningful.","PeriodicalId":223584,"journal":{"name":"Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation","volume":"113 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"43","resultStr":"{\"title\":\"How to Compare Summarizers without Target Length? Pitfalls, Solutions and Re-Examination of the Neural Summarization Literature\",\"authors\":\"Simeng Sun, Ori Shapira, Ido Dagan, A. Nenkova\",\"doi\":\"10.18653/v1/W19-2303\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"We show that plain ROUGE F1 scores are not ideal for comparing current neural systems which on average produce different lengths. This is due to a non-linear pattern between ROUGE F1 and summary length. To alleviate the effect of length during evaluation, we have proposed a new method which normalizes the ROUGE F1 scores of a system by that of a random system with same average output length. A pilot human evaluation has shown that humans prefer short summaries in terms of the verbosity of a summary but overall consider longer summaries to be of higher quality. While human evaluations are more expensive in time and resources, it is clear that normalization, such as the one we proposed for automatic evaluation, will make human evaluations more meaningful.\",\"PeriodicalId\":223584,\"journal\":{\"name\":\"Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation\",\"volume\":\"113 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"43\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.18653/v1/W19-2303\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/W19-2303","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 43

摘要

我们表明，普通的ROUGE F1分数对于比较平均产生不同长度的当前神经系统并不理想。这是由于ROUGE F1和总结长度之间的非线性模式。为了减轻长度在评价过程中的影响，我们提出了一种新的方法，即用具有相同平均输出长度的随机系统的ROUGE F1分数对系统的ROUGE F1分数进行归一化。一项初步的人类评估表明，就摘要的冗长程度而言，人类更喜欢简短的摘要，但总体而言，人们认为较长的摘要质量更高。虽然人类评估在时间和资源上更昂贵，但很明显，规范化，例如我们为自动评估提出的规范化，将使人类评估更有意义。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

How to Compare Summarizers without Target Length? Pitfalls, Solutions and Re-Examination of the Neural Summarization Literature

We show that plain ROUGE F1 scores are not ideal for comparing current neural systems which on average produce different lengths. This is due to a non-linear pattern between ROUGE F1 and summary length. To alleviate the effect of length during evaluation, we have proposed a new method which normalizes the ROUGE F1 scores of a system by that of a random system with same average output length. A pilot human evaluation has shown that humans prefer short summaries in terms of the verbosity of a summary but overall consider longer summaries to be of higher quality. While human evaluations are more expensive in time and resources, it is clear that normalization, such as the one we proposed for automatic evaluation, will make human evaluations more meaningful.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation

自引率

0.00%

发文量