Article
A Comparative Survey of CNN-LSTM Architectures for Image Captioning
2025-12-15
Abstract excerpt
Image captioning, the task of automatically generating textual descriptions for images, lies at the intersection of computer vision and natural language processing. Architectures combining Convolutional Neural Networks (CNNs) for visual feature extraction and Long Short-Term Memory (LSTM) networks for language generation have become a dominant paradigm. This survey provides a comprehensive overview of fifteen infl...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 7c9fe410-6a4e-5371-b51d-b976f6aee789
- DOI
- 10.20944/preprints202512.1301.v1
