Learning Textual Representations from Multiple Modalities to Detect Fake News Through One-Class Learning

Proceedings of the Brazilian Symposium on Multimedia and the Web Pub Date : 2021-09-27 DOI:10.1145/3470482.3479634

M. Gôlo, M. C. D. Souza, R. G. Rossi, S. O. Rezende, B. Nogueira, R. Marcacini

{"title":"Learning Textual Representations from Multiple Modalities to Detect Fake News Through One-Class Learning","authors":"M. Gôlo, M. C. D. Souza, R. G. Rossi, S. O. Rezende, B. Nogueira, R. Marcacini","doi":"10.1145/3470482.3479634","DOIUrl":null,"url":null,"abstract":"Fake news can rapidly spread through internet users. Approaches proposed in the literature for content classification usually learn models considering textual and contextual features from real and fake news to minimize the spread of disinformation. One of the prominent approaches to detect fake news is One-Class Learning (OCL), as it minimizes the data labeling effort, requiring only the labeling of fake news documents. The performance of these algorithms depends on the structured representation of the documents used in the learning process. Generally, a textual-based unimodal representation is used, such as bag-of-words or representations based on linguistic categories. We propose MVAE-FakeNews, a multimodal representation method to detect fake news in OCL. The proposed approach uses a Multimodal Variational Autoencoder, learns a new representation from the combination of two modalities considered promising for fake news detection: text embeddings and topic information. In the experiments, we used three datasets considering Portuguese and English languages. Results show that the MVAE-FakeNews obtained a better F1-Score for the class of interest, outperforming another nine methods in ten of twelve evaluated scenarios. MVAE-FakeNews presented a better average ranking and statistical difference from other representation models. The proposed method proved to be promising to represent the texts in the OCL scenario to detect fake news.","PeriodicalId":350776,"journal":{"name":"Proceedings of the Brazilian Symposium on Multimedia and the Web","volume":"83 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2021-09-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"7","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the Brazilian Symposium on Multimedia and the Web","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3470482.3479634","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 7

Abstract

Fake news can rapidly spread through internet users. Approaches proposed in the literature for content classification usually learn models considering textual and contextual features from real and fake news to minimize the spread of disinformation. One of the prominent approaches to detect fake news is One-Class Learning (OCL), as it minimizes the data labeling effort, requiring only the labeling of fake news documents. The performance of these algorithms depends on the structured representation of the documents used in the learning process. Generally, a textual-based unimodal representation is used, such as bag-of-words or representations based on linguistic categories. We propose MVAE-FakeNews, a multimodal representation method to detect fake news in OCL. The proposed approach uses a Multimodal Variational Autoencoder, learns a new representation from the combination of two modalities considered promising for fake news detection: text embeddings and topic information. In the experiments, we used three datasets considering Portuguese and English languages. Results show that the MVAE-FakeNews obtained a better F1-Score for the class of interest, outperforming another nine methods in ten of twelve evaluated scenarios. MVAE-FakeNews presented a better average ranking and statistical difference from other representation models. The proposed method proved to be promising to represent the texts in the OCL scenario to detect fake news.

查看原文本刊更多论文

通过一堂课学习，学习多种形式的文本表示来检测假新闻

假新闻可以通过互联网用户迅速传播。文献中提出的内容分类方法通常学习考虑真假新闻文本和上下文特征的模型，以最大限度地减少虚假信息的传播。检测假新闻的主要方法之一是单类学习(OCL)，因为它最大限度地减少了数据标记工作，只需要标记假新闻文档。这些算法的性能取决于学习过程中使用的文档的结构化表示。一般使用基于文本的单模态表示，如词袋表示或基于语言类别的表示。我们提出了一种多模态表示方法mvee - fakenews来检测OCL中的假新闻。提出的方法使用多模态变分自编码器，从文本嵌入和主题信息这两种被认为有希望用于假新闻检测的模态组合中学习新的表示。在实验中，我们使用了三个考虑葡萄牙语和英语语言的数据集。结果表明，MVAE-FakeNews在感兴趣的类别中获得了更好的f1分，在12个评估场景中的10个中优于其他9个方法。与其他表征模型相比，MVAE-FakeNews表现出更好的平均排名和统计差异。所提出的方法被证明有希望表示OCL场景中的文本来检测假新闻。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

Proceedings of the Brazilian Symposium on Multimedia and the Web

自引率

0.00%

发文量