Back to search

Article

Unified Generative Vision-Language Understanding

2024-11-29

Abstract excerpt

This paper introduces an innovative learning framework where linguistic representations are inherently grounded in visual perceptions, circumventing the need for predefined categorical structures. The proposed method, termed Generative Semantic Embedding Model (GSEM), employs a unified generative strategy to construct a shared semantic-visual embedding space. This embedding facilitates robust language grounding ac...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
5070e8c7-1421-5096-9a38-9d2437b8bb80
DOI
10.20944/preprints202411.2316.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Unified Generative Vision-Language UnderstandingDOI 10.20944/preprints202411.2316.v1
Select a neighboring publication to make it the new centre.