Back to search

Article

Expressive Speech Synthesis by Modeling Prosody with Variational Autoencoders for Bangla Text-to-Speech

2022-06-08

Abstract excerpt

With the advent of deep learning, Text-to-Speech (TTS) research has made a great leap in producing natural speech. The state-of-the-art TTS systems generate average prosody, resulting in a lack of variety and expressiveness found in human speech. To avoid synthesizing monotonous speech and averaged prosody, it is desirable to have a way of modeling the variation in the speech prosody. To generate highly expressive...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
9f207b31-2006-5c11-af57-58710ff775e4
DOI
10.21203/rs.3.rs-1690533/v2
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Expressive Speech Synthesis by Modeling Prosody with Variational Autoencoders for Bangla Text-to-SpeechDOI 10.21203/rs.3.rs-1690533/v2
Select a neighboring publication to make it the new centre.