Back to search

Article

Enhancing Caption Fidelity via Explanation-Guided Captioning with Vision-Language Fine-Tuning

2025-08-01

Abstract excerpt

Image captioning models have achieved remarkable progress with the introduction of attention mechanisms and transformer-based architectures. However, understanding and diagnosing their predictions remain a challenging task, particularly in terms of attribution, interpretability, and mitigation of hallucinated outputs. In this work, we present \textbf{CAPEV}, a novel explanation-guided fine-tuning paradigm that bui...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
9d8196b4-3ae3-59ab-bf5b-1fd907be1821
DOI
10.20944/preprints202508.0076.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Enhancing Caption Fidelity via Explanation-Guided Captioning with Vision-Language Fine-TuningDOI 10.20944/preprints202508.0076.v1
Select a neighboring publication to make it the new centre.