Back to search

Article

A Review of Multimodal Vision-Language Models: Foundations, Applications, and Future Directions

2025-10-31

Abstract excerpt

Large Language Models (LLMs) have rapidly become a central focus in both research and practical applications, owing to their remarkable ability to understand and generate text with a level of fluency comparable to human communication. Recently, these models have evolved into multimodal large language models (MM-LLMs), extending their capabilities beyond text to include images, audio, and video. This advancement ha...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
60060811-eea7-5290-b512-9bae058cc596
DOI
10.22541/au.176184212.29642407/v2
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A Review of Multimodal Vision-Language Models: Foundations, Applications, and Future DirectionsDOI 10.22541/au.176184212.29642407/v2
Select a neighboring publication to make it the new centre.