Article
FineRegion-LM: Enhancing Large Vision-Language Models for Fine-Grained Region-Level Understanding
2024-12-16
Abstract excerpt
Large Vision-Language Models (LVLMs) have achieved remarkable success in vision-language tasks, yet they often fall short in fine-grained region-level understanding due to limited spatial sensitivity and insufficient region-specific annotations. To address these challenges, we propose FineRegion-LM, a generative model that enhances LVLMs' capabilities in region comprehension through a novel dual-stage framework. O...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f839f4b8-0c68-5209-ad44-a2c9c0d4619a
- DOI
- 10.20944/preprints202412.1262.v1
