Article
A Review of Multimodal Vision–Language Models: Foundations, Applications, and Future Directions
2025-11-03
Abstract excerpt
Large Language Models (LLMs) have rapidly become a central focus in both research and practical applications, owing to their remarkable ability to understand and generate text with a level of fluency comparable to human communication. Recently, these models have evolved into multimodal large language models (MM-LLMs), extending their capabilities beyond text to include images, audio, and video. This advancement ha...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 1e4fe3db-c98c-5921-a1db-ced83a991590
- DOI
- 10.20944/preprints202510.2511.v1
