FINE-GRAINED EMOTIONAL CONTROL OF TEXT-TO-SPEECH: LEARNING TO RANK INTER- AND INTRA-CLASS EMOTION INTENSITIES

Abstract: State-of-the-art Text-To-Speech (TTS) models are capable of producing high-quality speech. The generated speech, however, is usually neutral in emotional expression, whereas very often one would want fine-grained emotional control of words or phonemes. Although still challenging, the first TTS models have been recently proposed that are able to control voice by manually assigning emotion intensity. Unfortunately, due to the neglect of intra-class distance, the intensity differences are often unrecognizable. In this paper, we propose a fine-grained controllable emotional TTS, that considers both inter- and intra-class distances and be able to synthesize speech with recognizable intensity difference. Our subjective and objective experiments demonstrate that our model exceeds two state-of-the-art controllable TTS models for controllability, emotion expressiveness and naturalness.

Emotion Intensity Control

We use min, median and max intensity representations, respectively, as conditions for each phoneme to synthesize speech. For "increase", we gradually increase the intensity, from min to median, then to max. For "decrease", it's the other way around.

Text: From that moment, his friendship for Belize turns to hatred and jealousy.

Speaker: Jenie | Emotion: Neutral

Neutral Amused Disgusted Angry Sleepy
Speaker Jenie Bea Sam Josh Jenie Bea Sam Josh Jenie Bea Sam Josh* Jenie Bea Sam Josh* Jenie Bea Sam Josh
Intensity
/
/
/
/
/
/
/
/
/
/

*In the dataset, Josh has no data for Disgusted and Angry.

 

Intensity Controllability Comparison

We compare our model with two baselines, regarding the difference between synthesized speech with Min, Median and Max intensity representations.

RFTacotron: Lee, Younggun and Taesu Kim. "Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis." ICASSP 2019.

FEC: Lei, Yinjiao et al. "Fine-Grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis." SLT 2021.
RFTacotron FEC Ours
Emotion: Amused   Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Emotion: Angry   Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Emotion: Disgusted   Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Emotion: Sleepy   Min:  

Median:  

Max:  
Min:  

Median:  

Max:  
Min:  

Median:  

Max:  

Naturnalness Comparison

We compare the naturalness of synthesized speech. For FEC and Our model, the intensity for each phoneme is assigned by their individual Rank models. For RFTacotron, we use ground truth samples as references.

RFTacotron: Lee, Younggun and Taesu Kim. "Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis." ICASSP 2019.

FEC: Lei, Yinjiao et al. "Fine-Grained Emotion Strength Transfer, Control and Prediction for Emotional Speech Synthesis." SLT 2021.
Ground Truth RFTacotron FEC Ours