A Survey on VQA: Datasets and Approaches

2020 2nd International Conference on Information Technology and Computer Application (ITCA) Pub Date : 2020-12-01 DOI:10.1109/ITCA52113.2020.00069

Yeyun Zou, Qiyu Xie

引用次数: 7

Abstract

Visual question answering (VQA) is a task that combines both the techniques of computer vision and natural language processing. It requires models to answer a text-based question according to the information contained in a visual. In recent years, the research field of VQA has been expanded. Research that focuses on the VQA, examining the reasoning ability and VQA on scientific diagrams, has also been explored more. Meanwhile, more multimodal feature fusion mechanisms have been proposed. This paper will review and analyze existing datasets, metrics, and models proposed for the VQA task.

查看原文本刊更多论文

VQA调查:数据集和方法

视觉问答(VQA)是计算机视觉和自然语言处理技术相结合的一项任务。它要求模型根据图像中包含的信息回答基于文本的问题。近年来，VQA的研究领域不断扩大。研究的重点是VQA，考察科学图表的推理能力和VQA，也得到了更多的探索。同时，还提出了更多的多模态特征融合机制。本文将回顾和分析为VQA任务提出的现有数据集、度量和模型。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

2020 2nd International Conference on Information Technology and Computer Application (ITCA)

自引率

0.00%

发文量