Back to search

Article

A Review of Multimodal Vision–Language Models: Foundations, Applications, and Future Directions

2025-11-03

Abstract excerpt

Large Language Models (LLMs) have rapidly become a central focus in both research and practical applications, owing to their remarkable ability to understand and generate text with a level of fluency comparable to human communication. Recently, these models have evolved into multimodal large language models (MM-LLMs), extending their capabilities beyond text to include images, audio, and video. This advancement ha...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
1e4fe3db-c98c-5921-a1db-ced83a991590
DOI
10.20944/preprints202510.2511.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A Review of Multimodal Vision–Language Models: Foundations, Applications, and Future DirectionsDOI 10.20944/preprints202510.2511.v1
Select a neighboring publication to make it the new centre.