现有标注语料库中攻击性语言类别的语义分析

Uporabna informatika Pub Date : 2022-05-04 DOI:10.31449/upinf.vol30.num1.151

Maša Kljun, Matija Teršek, Slavko Žitnik

{"title":"现有标注语料库中攻击性语言类别的语义分析","authors":"Maša Kljun, Matija Teršek, Slavko Žitnik","doi":"10.31449/upinf.vol30.num1.151","DOIUrl":null,"url":null,"abstract":"\nThere exists a vast amount of different offensive language corpora for English language, annotation criteria and category naming. In this paper, we explore 21 different categories of offensive language. We use natural language processing techniques to find correlations between the categories based on seven different data sets. We employ several traditional (TF–IDF) and advanced (fastText, GloVe, Word2Vec, BERT, and other deep NLP methods) techniques to uncover similarities among different offensive language categories. The findings reveal that most of the categories are densely interconnected, while a two-level hierarchical representation of them can be provided. We also transfer the analysis to the Slovenian language and compare the findings between both researched languages.\n","PeriodicalId":393713,"journal":{"name":"Uporabna informatika","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-05-04","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Semantic analysis of offensive language categories from existing annotated corpora\",\"authors\":\"Maša Kljun, Matija Teršek, Slavko Žitnik\",\"doi\":\"10.31449/upinf.vol30.num1.151\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"\\nThere exists a vast amount of different offensive language corpora for English language, annotation criteria and category naming. In this paper, we explore 21 different categories of offensive language. We use natural language processing techniques to find correlations between the categories based on seven different data sets. We employ several traditional (TF–IDF) and advanced (fastText, GloVe, Word2Vec, BERT, and other deep NLP methods) techniques to uncover similarities among different offensive language categories. The findings reveal that most of the categories are densely interconnected, while a two-level hierarchical representation of them can be provided. We also transfer the analysis to the Slovenian language and compare the findings between both researched languages.\\n\",\"PeriodicalId\":393713,\"journal\":{\"name\":\"Uporabna informatika\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-05-04\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Uporabna informatika\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.31449/upinf.vol30.num1.151\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Uporabna informatika","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.31449/upinf.vol30.num1.151","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

英语语言中存在着大量不同的攻击性语言语料库、标注标准和类别命名。在本文中，我们探讨了21种不同类别的攻击性语言。我们使用自然语言处理技术来发现基于七个不同数据集的类别之间的相关性。我们使用了几种传统的(TF-IDF)和高级的(fastText, GloVe, Word2Vec, BERT和其他深度NLP方法)技术来发现不同攻击性语言类别之间的相似性。研究结果表明，大多数类别是紧密相连的，而它们的两级层次表示可以提供。我们还将分析转移到斯洛文尼亚语，并比较两种语言之间的研究结果。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

Semantic analysis of offensive language categories from existing annotated corpora

There exists a vast amount of different offensive language corpora for English language, annotation criteria and category naming. In this paper, we explore 21 different categories of offensive language. We use natural language processing techniques to find correlations between the categories based on seven different data sets. We employ several traditional (TF–IDF) and advanced (fastText, GloVe, Word2Vec, BERT, and other deep NLP methods) techniques to uncover similarities among different offensive language categories. The findings reveal that most of the categories are densely interconnected, while a two-level hierarchical representation of them can be provided. We also transfer the analysis to the Slovenian language and compare the findings between both researched languages.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Uporabna informatika

自引率

0.00%

发文量