一种用于视觉问答的增强词加权问题嵌入

J. Inf. Knowl. Manag. Pub Date : 2022-05-05 DOI:10.1142/s0219649222500289

Sruthy Manmadhan, Binsu C. Kovoor

{"title":"一种用于视觉问答的增强词加权问题嵌入","authors":"Sruthy Manmadhan, Binsu C. Kovoor","doi":"10.1142/s0219649222500289","DOIUrl":null,"url":null,"abstract":"Visual Question Answering (VQA) is a multi-modal AI-complete task of answering natural language questions about images. Literature solved VQA with a three-phase pipeline: image and question featurisation, multi-modal feature fusion and answer generation or prediction. Most of the works have given attention to the second phase, where multi-modal features get combined ignoring the effect of individual input features. This work investigates VQA’s natural language question embedding phase by proposing a new question featurisation framework based on Supervised Term Weighting (STW) schemes. In addition, two new STW schemes integrating text semantics, qf.cos and tf.rf.sim, have been introduced to boost the framework’s performance. A series of tests on the DAQUAR VQA dataset is used to compare the new system to conventional pre-trained word embedding. Over the past few years, STW schemes have been commonly used in text classification research. In light of this, tests are carried out to verify the effectiveness of the two newly proposed STW schemes in the general text classification task.","PeriodicalId":127309,"journal":{"name":"J. Inf. Knowl. Manag.","volume":"12 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-05-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"An Enhanced Term Weighted Question Embedding for Visual Question Answering\",\"authors\":\"Sruthy Manmadhan, Binsu C. Kovoor\",\"doi\":\"10.1142/s0219649222500289\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Visual Question Answering (VQA) is a multi-modal AI-complete task of answering natural language questions about images. Literature solved VQA with a three-phase pipeline: image and question featurisation, multi-modal feature fusion and answer generation or prediction. Most of the works have given attention to the second phase, where multi-modal features get combined ignoring the effect of individual input features. This work investigates VQA’s natural language question embedding phase by proposing a new question featurisation framework based on Supervised Term Weighting (STW) schemes. In addition, two new STW schemes integrating text semantics, qf.cos and tf.rf.sim, have been introduced to boost the framework’s performance. A series of tests on the DAQUAR VQA dataset is used to compare the new system to conventional pre-trained word embedding. Over the past few years, STW schemes have been commonly used in text classification research. In light of this, tests are carried out to verify the effectiveness of the two newly proposed STW schemes in the general text classification task.\",\"PeriodicalId\":127309,\"journal\":{\"name\":\"J. Inf. Knowl. Manag.\",\"volume\":\"12 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2022-05-05\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"J. Inf. Knowl. Manag.\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1142/s0219649222500289\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"J. Inf. Knowl. Manag.","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1142/s0219649222500289","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

摘要

视觉问答(VQA)是一种多模式的人工智能完成任务，用于回答有关图像的自然语言问题。文献用一个三相管道来解决VQA:图像和问题特征化、多模态特征融合和答案生成或预测。大多数研究都关注第二阶段，即多模态特征的组合，忽略了单个输入特征的影响。本文通过提出一种基于监督项加权(STW)方案的问题特征化框架，研究了VQA的自然语言问题嵌入阶段。此外，还提出了两种新的集成文本语义的STW方案qf。Cos和tf。rf。已经引入了Sim，以提高框架的性能。在DAQUAR VQA数据集上进行了一系列测试，将新系统与传统的预训练词嵌入进行了比较。在过去的几年里，STW算法被广泛应用于文本分类研究。基于此，本文通过实验验证了两种新提出的STW方案在一般文本分类任务中的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文本刊更多论文

An Enhanced Term Weighted Question Embedding for Visual Question Answering

Visual Question Answering (VQA) is a multi-modal AI-complete task of answering natural language questions about images. Literature solved VQA with a three-phase pipeline: image and question featurisation, multi-modal feature fusion and answer generation or prediction. Most of the works have given attention to the second phase, where multi-modal features get combined ignoring the effect of individual input features. This work investigates VQA’s natural language question embedding phase by proposing a new question featurisation framework based on Supervised Term Weighting (STW) schemes. In addition, two new STW schemes integrating text semantics, qf.cos and tf.rf.sim, have been introduced to boost the framework’s performance. A series of tests on the DAQUAR VQA dataset is used to compare the new system to conventional pre-trained word embedding. Over the past few years, STW schemes have been commonly used in text classification research. In light of this, tests are carried out to verify the effectiveness of the two newly proposed STW schemes in the general text classification task.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

J. Inf. Knowl. Manag.

自引率

0.00%

发文量