Back to search

Article

Towards Human-Centered and Efficient Video Synthesis: A Survey of Multimodal Diffusion Models

2025-10-07

Abstract excerpt

<title>Abstract</title> <p>Multimodal video diffusion models have emerged as transformative tools for controlled video synthesis, integrating text, images, audio, and pose sequences to generate semantically meaningful content. Despite significant advances, critical gaps persist in temporal consistency, multimodal alignment, and human-centric motion generation. Existing surveys have not addressed clearly the compl...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
932cda94-5921-5433-a4cd-ed9a491129e6
DOI
10.21203/rs.3.rs-7533477/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Towards Human-Centered and Efficient Video Synthesis: A Survey of Multimodal Diffusion ModelsDOI 10.21203/rs.3.rs-7533477/v1
Select a neighboring publication to make it the new centre.