Article
Expressive Speech Synthesis by Modeling Prosody with Variational Autoencoders for Bangla Text-to-Speech
2022-06-08
Abstract excerpt
With the advent of deep learning, Text-to-Speech (TTS) research has made a great leap in producing natural speech. The state-of-the-art TTS systems generate average prosody, resulting in a lack of variety and expressiveness found in human speech. To avoid synthesizing monotonous speech and averaged prosody, it is desirable to have a way of modeling the variation in the speech prosody. To generate highly expressive...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 9f207b31-2006-5c11-af57-58710ff775e4
- DOI
- 10.21203/rs.3.rs-1690533/v2
