Back to search

Article

Fusing Text-Speech Multimodal Cues with 2D Motion Priors: Boosting 3D Human Motion Generation for VR and Animation

2025-12-02

Abstract excerpt

The fields of virtual reality (VR) and animation require 3D human motion that matches both text semantics and speech rhythm, yet existing approaches (such as T3M proposed by Peng et al., 2024) suffer from insufficient support of high-quality 3D data. To address this issue, this paper proposes a novel solution: integrating text-speech multimodal fusion with 2D motion priors (adopted from Motion-2-to-3 by Pi et al.,...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
8bbd5b9b-a7ef-5173-b81d-5a4b62f64366
DOI
10.22541/au.176463747.73317510/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Fusing Text-Speech Multimodal Cues with 2D Motion Priors: Boosting 3D Human Motion Generation for VR and AnimationDOI 10.22541/au.176463747.73317510/v1
Select a neighboring publication to make it the new centre.