Article
Fusing Text-Speech Multimodal Cues with 2D Motion Priors: Boosting 3D Human Motion Generation for VR and Animation
2025-12-02
Abstract excerpt
The fields of virtual reality (VR) and animation require 3D human motion that matches both text semantics and speech rhythm, yet existing approaches (such as T3M proposed by Peng et al., 2024) suffer from insufficient support of high-quality 3D data. To address this issue, this paper proposes a novel solution: integrating text-speech multimodal fusion with 2D motion priors (adopted from Motion-2-to-3 by Pi et al.,...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 8bbd5b9b-a7ef-5173-b81d-5a4b62f64366
- DOI
- 10.22541/au.176463747.73317510/v1
