Article
Enhancing Caption Fidelity via Explanation-Guided Captioning with Vision-Language Fine-Tuning
2025-08-01
Abstract excerpt
Image captioning models have achieved remarkable progress with the introduction of attention mechanisms and transformer-based architectures. However, understanding and diagnosing their predictions remain a challenging task, particularly in terms of attribution, interpretability, and mitigation of hallucinated outputs. In this work, we present \textbf{CAPEV}, a novel explanation-guided fine-tuning paradigm that bui...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9d8196b4-3ae3-59ab-bf5b-1fd907be1821
- DOI
- 10.20944/preprints202508.0076.v1
