{"title":"多文档摘要中的子主题句子评分","authors":"Sujian Li, Weiguang Qu","doi":"10.1109/ALPIT.2007.106","DOIUrl":null,"url":null,"abstract":"In previous works, subtopics are seldom mentioned in multi-document summarization while only one topic is focused to extract summary. In this paper, we propose a subtopic- focused model to score sentences in the extractive summarization task. Different with supervised methods, it does not require costly manual work to form the training set. Multiple documents are represented as mixture over subtopics, denoted by term distributions through unsupervised learning. Our method learns the subtopic distribution over sentences via a hierarchical Bayesian model, through which sentences are scored and extracted as summary. Experiments on DUC 2006 data are performed and the ROUGE evaluation results show that the proposed method can reach the state-of-the-art performance.","PeriodicalId":412433,"journal":{"name":"Advanced Language Processing and Web Information Technology","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"1900-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Subtopic-Focused Sentence Scoring in Multi-document Summarization\",\"authors\":\"Sujian Li, Weiguang Qu\",\"doi\":\"10.1109/ALPIT.2007.106\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In previous works, subtopics are seldom mentioned in multi-document summarization while only one topic is focused to extract summary. In this paper, we propose a subtopic- focused model to score sentences in the extractive summarization task. Different with supervised methods, it does not require costly manual work to form the training set. Multiple documents are represented as mixture over subtopics, denoted by term distributions through unsupervised learning. Our method learns the subtopic distribution over sentences via a hierarchical Bayesian model, through which sentences are scored and extracted as summary. Experiments on DUC 2006 data are performed and the ROUGE evaluation results show that the proposed method can reach the state-of-the-art performance.\",\"PeriodicalId\":412433,\"journal\":{\"name\":\"Advanced Language Processing and Web Information Technology\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"1900-01-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Advanced Language Processing and Web Information Technology\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ALPIT.2007.106\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Advanced Language Processing and Web Information Technology","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ALPIT.2007.106","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
Subtopic-Focused Sentence Scoring in Multi-document Summarization
In previous works, subtopics are seldom mentioned in multi-document summarization while only one topic is focused to extract summary. In this paper, we propose a subtopic- focused model to score sentences in the extractive summarization task. Different with supervised methods, it does not require costly manual work to form the training set. Multiple documents are represented as mixture over subtopics, denoted by term distributions through unsupervised learning. Our method learns the subtopic distribution over sentences via a hierarchical Bayesian model, through which sentences are scored and extracted as summary. Experiments on DUC 2006 data are performed and the ROUGE evaluation results show that the proposed method can reach the state-of-the-art performance.