Article
RAE-NeRF: Residual-Based Audio-Video Encoder with Denoising in Talking Head Synchronization
2025-09-26
Abstract excerpt
In recent years, speech-driven facial synthesis has attracted significant attention due to its wide applications in virtual humans, remote conferencing, and digital human generation. However, existing methods still face limitations in terms of realism, synchronization, and robustness, primarily due to noise interference in speech signals and insufficient precision in audio-visual feature fusion. To address these c...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 5cf39b87-f231-5f78-95d2-7fadea7ad6a4
- DOI
- 10.20944/preprints202509.2231.v1
