基于K-MHaS的韩语在线新闻评论多标签仇恨言论检测数据集

Proceedings of COLING. International Conference on Computational Linguistics Pub Date : 2022-08-23 DOI:10.48550/arXiv.2208.10684

Jean Lee, Taejun Lim, Hee-Youn Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, S. Han

{"title":"基于K-MHaS的韩语在线新闻评论多标签仇恨言论检测数据集","authors":"Jean Lee, Taejun Lim, Hee-Youn Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, S. Han","doi":"10.48550/arXiv.2208.10684","DOIUrl":null,"url":null,"abstract":"Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baselines on K-MHaS. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class.","PeriodicalId":91381,"journal":{"name":"Proceedings of COLING. International Conference on Computational Linguistics","volume":"29 1","pages":"3530-3538"},"PeriodicalIF":0.0000,"publicationDate":"2022-08-23","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"5","resultStr":"{\"title\":\"K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment\",\"authors\":\"Jean Lee, Taejun Lim, Hee-Youn Lee, Bogeun Jo, Yangsok Kim, Heegeun Yoon, S. Han\",\"doi\":\"10.48550/arXiv.2208.10684\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baselines on K-MHaS. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class.\",\"PeriodicalId\":91381,\"journal\":{\"name\":\"Proceedings of COLING. International Conference on Computational Linguistics\",\"volume\":\"29 1\",\"pages\":\"3530-3538\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-08-23\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"5\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Proceedings of COLING. International Conference on Computational Linguistics\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.48550/arXiv.2208.10684\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of COLING. International Conference on Computational Linguistics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.48550/arXiv.2208.10684","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 5

摘要

由于在线内容的增长，在线仇恨言论检测已成为一个重要问题，但英语以外的语言资源极其有限。我们引入了K-MHaS，这是一种新的多标签数据集，用于仇恨言论检测，可以有效地处理韩语模式。该数据集由来自新闻评论的109k个话语组成，使用1到4个标签进行多标签分类，并处理主观性和交叉性。我们评估K-MHaS的强基线。带有子字符标记器的KR-BERT优于其他工具，可以识别每个仇恨言论类别中的分解字符。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

K-MHaS: A Multi-label Hate Speech Detection Dataset in Korean Online News Comment

Online hate speech detection has become an important issue due to the growth of online content, but resources in languages other than English are extremely limited. We introduce K-MHaS, a new multi-label dataset for hate speech detection that effectively handles Korean language patterns. The dataset consists of 109k utterances from news comments and provides a multi-label classification using 1 to 4 labels, and handles subjectivity and intersectionality. We evaluate strong baselines on K-MHaS. KR-BERT with a sub-character tokenizer outperforms others, recognizing decomposed characters in each hate speech class.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Proceedings of COLING. International Conference on Computational Linguistics

自引率

0.00%

发文量