Visual Question Answering
视觉问答VQACommonGiven an image and a natural-language question about it, the model answers in words.
Visual question answering takes an image and a question about it, such as “how many cups are on the table?”, and outputs a natural-language answer, requiring the model to understand the image, the language, and relevant common sense together. The task was introduced by Antol, Agrawal, and colleagues at ICCV 2015; the VQA dataset has about 250,000 images and 760,000 questions. The 2017 VQA v2 rebalanced the dataset to reduce cases where a model could guess the answer from the question alone, without looking at the image. Today most vision-language models are trained and evaluated with question-answering formats, and embodied and spatial-reasoning benchmarks like ERQA and VSI-Bench are also framed as question answering; VQA data is often mixed into VLA training as a co-training task — for example, π0.5 uses web data including VQAv2 to preserve its image-understanding ability. Unlike embodied question answering, VQA only looks at a given image and never requires the robot to move around to find the answer.
ExampleGiven a photo of a kitchen counter, ask “is the cup to the left of the sink empty?” and the model answers “yes, it’s empty”; an embodied-reasoning benchmark would instead ask something like “which object should the robot’s gripper move above first?”
- Also called
- VQA, Image QA
- Related
- Vision-Language Model · Multimodal Large Language Model · Embodied Question Answering · Co-training · ERQA · VSI-Bench
- Sources
- VQA: Visual Question Answering (ICCV 2015)
VQA 官网(VQA v2 数据集与挑战赛) (Chinese)
π0.5: a Vision-Language-Action Model with Open-World Generalization(协同训练用到 VQAv2 等网络数据) (Chinese) - As of
- 2025-04