Article
Towards Human-Centered and Efficient Video Synthesis: A Survey of Multimodal Diffusion Models
2025-10-07
Abstract excerpt
<title>Abstract</title> <p>Multimodal video diffusion models have emerged as transformative tools for controlled video synthesis, integrating text, images, audio, and pose sequences to generate semantically meaningful content. Despite significant advances, critical gaps persist in temporal consistency, multimodal alignment, and human-centric motion generation. Existing surveys have not addressed clearly the compl...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 932cda94-5921-5433-a4cd-ed9a491129e6
- DOI
- 10.21203/rs.3.rs-7533477/v1
