Back to search

Article

A Unified Framework for Human Motion Generation with Multimodal Inputs

2025-08-28

Abstract excerpt

<title>Abstract</title> <p>To enable generalized human motion generation, this paper proposes a unified generation framework, UniMotion, which supports multimodal inputs including text, image and audio. The method uses a unified prompt encoder to map different inputs into a shared cross-modal semantic space. It adopts a two-stage motion decoder to gradually generate fine-grained skeleton sequences. A multimodal a...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
f2fd87f4-4d00-573b-be42-02f45cb9fdcb
DOI
10.21203/rs.3.rs-7467386/v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
A Unified Framework for Human Motion Generation with Multimodal InputsDOI 10.21203/rs.3.rs-7467386/v1
Select a neighboring publication to make it the new centre.