Article
Cognitive Grounding for Visual Question Reasoning via Dynamic Knowledge Imagination
2025-10-27
Abstract excerpt
Visual question answering (VQA) represents a critical intersection of vision and language understanding, where models must perceive visual scenes and reason about their underlying semantics. However, human-like reasoning often extends beyond what is directly observable—requiring the invocation of prior knowledge, inference, and commonsense understanding. In this work, we reexamine the nature of external knowledge...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- c441130e-4649-5e05-8a6f-e61f006db1ad
- DOI
- 10.20944/preprints202510.1967.v1
