Article
A Unified Framework for Human Motion Generation with Multimodal Inputs
2025-08-28
Abstract excerpt
<title>Abstract</title> <p>To enable generalized human motion generation, this paper proposes a unified generation framework, UniMotion, which supports multimodal inputs including text, image and audio. The method uses a unified prompt encoder to map different inputs into a shared cross-modal semantic space. It adopts a two-stage motion decoder to gradually generate fine-grained skeleton sequences. A multimodal a...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- f2fd87f4-4d00-573b-be42-02f45cb9fdcb
- DOI
- 10.21203/rs.3.rs-7467386/v1
