Article
Unified Generative Vision-Language Understanding
2024-11-29
Abstract excerpt
This paper introduces an innovative learning framework where linguistic representations are inherently grounded in visual perceptions, circumventing the need for predefined categorical structures. The proposed method, termed Generative Semantic Embedding Model (GSEM), employs a unified generative strategy to construct a shared semantic-visual embedding space. This embedding facilitates robust language grounding ac...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 5070e8c7-1421-5096-9a38-9d2437b8bb80
- DOI
- 10.20944/preprints202411.2316.v1
