HANOI UNIVERSITY OF SCIENCE AND TECHNOLOGY MASTER THESIS Expressive Speech Synthesis NGUYEN THI NGOC ANH Anh.vn School of Information and Communication Technology Supervisor: Dr. Nguyen Thanh Hung Supervisor’s signature School: Information and Communication Technology 18th May 2023 Graduation Thesis Assignment Name: Nguyen Thi Ngoc Anh Phone: +84342612379 Email: Anh.vn; ngocanh2162@gmail.com Class: CH2021A Affiliation: Hanoi University of Science and Technology Nguyen Thi Ngoc Anh - hereby warrants that the work and presentation in this thesis were performed by myself under the supervision of Dr. Nguyen Thanh Hung. All the results presented in this thesis are truthful and are not copied from any other works.
All references in this thesis including images, tables, figures, and quotes are clearly and fully documented in the bibliography. I will take full responsibility for even one copy that violates school regulations. Student Signature and Name Acknowledgement I’d like to take this opportunity to thank everyone who has been so supportive of me throughout my academic career. To begin, I’d like to thank Dr.
Nguyen Thanh Hung for his unwavering support and encouragement throughout my mas- ter’s studies. His support and guidance have been instrumental in helping me achieve my academic goals. In addition, I’d like to thank Dr. Nguyen Thi Thu Trang and her colleagues in Lab 914 for their assistance in completing the experiments.
Their willingness to share their knowledge and skills has been very helpful to me, and I’ve learned a lot from them. Her knowledge, advice, and support were very important to my academic career, and I will always be grateful to her. I also want to thank Dr. Do Van Hai and my other coworkers at Viettel Cyberspace Center for their constant help and support during my master’s studies.
Their will- ingness to lend a hand and help me out when I needed it has been very important to me. I would not have been able to achieve academic success without their as- sistance. Aside from my academic mentors, I’m thankful to my family and friends for their constant support and encouragement. Their never-ending love and support have given me strength and pushed me to do well in school.
Finally, I would like to thank myself for persevering and not giving up. The jour- ney was difficult, but I am proud of myself for overcoming the challenges and reaching my academic goals. Abstract Text-to-speech technology, also known as TTS, is a type of assistive technol- ogy that converts written text into spoken words. The overall goal of the speech synthesis research community is to create natural sounding synthetic speech.
Cur- rently, there are many speech synthesis engines available on the market, each with its own strengths and weaknesses. Some engines focus on generating natural- sounding speech, while others focus on generating expressive speech. To increase naturalness, researchers have recently identified synthesizing emotional speech as a major research focus for the speech community. Expressive speech synthesis is the ability to convey emotions and attitudes through synthesized speech.
This is achieved by adding prosodic features like intonation, stress, and rhythm to the speech waveform. Vietnamese expressive speech research is scarce, to my knowl- edge. No datasets from these articles have been released. However, significant work remains in this field.
A large, high-quality dataset is needed to investigate Vietnamese expressive speech. This thesis (1) publishes two Vietnamese emotional speech datasets, (2) proposes a method for automatically building data, and (3) develops a model for synthe- sizing emotional speech. The proposed method for automatically building data helps reduce costs and time by extracting and labeling data from available data sources. Simultaneously, the applicability of the presented data is illustrated using the proposed emotional speech synthesis model.
Keywords: Speech Synthesis, Text To Speech, Expressive Speech Synthesis, Cor- pus Building Student Signature and Name TABLE OF CONTENTS INTRODUCTION.1 Non-emotional Features.2 Traditional Speech Synthesis Techniques.3 Modern Speech Synthesis Techniques .3 Expressive Speech Synthesis. BUILDING VIETNAMESE EMOTIONAL SPEECH DATASET 13 2.1 Existing Emotion Datasets .2 Data Processing Techniques .2 Pipeline For Building Emotional Speech Dataset .2 Target Speech Segmentation.1 Analysis of Pipeline Errors. EMOTIONAL SPEECH SYNTHESIS SYSTEM.1 Baseline Acoustic Model .2 Proposed Acoustic Model .3 Result and Discussion. 52 LIST OF FIGURES 1.1 An example of waveform, spectrogram, and mel-spectrogram.3 An example of modern TTS architecture.4 Typical acoustic models.5 Some expressive speech synthesis techniques.1 Pipeline for building an emotional speech dataset.2 Audio post-processing.4 F0 means in the TTH and LMH datasets.5 t-SNE visualizations of emotion embeddings in the TTH dataset.6 t-SNE visualizations of emotion embeddings in the LMH dataset.1 Pipeline for training baseline acoustic model.2 Baseline acoustic model architecture.3 Proposed acoustic model architecture.4 Detail of Emotion Encoder module.6 HifiGAN model architecture [66].
39 LIST OF TABLES 2.1 Some emotional datasets .2 Pipeline errors in the LMH dataset .3 LMH dataset before and after normalization .4 Compare manual pipeline and proposed pipeline processing times 27 2.5 Syllable coverage in two datasets .2 Acoustic model configuration .3 MOS score of data evaluation .4 EIR score of data evaluation .5 SUS score of data evaluation .6 MOS score in model evaluation .7 EIR score in model evaluation .8 SUS score in model evaluation. 50 ACRONYMS Notation Description AI Artificial Intelligence ASR Automatic Speech Recognition CNN Convolutional Neural Network DNN Deep Neural Network E2E End To End EIR Emotion Identification Rate ESS Expressive Speech Synthesis GAN Generative Adversarial Network GST Global Style Tokens HMM Hidden Markov Model LMH Luong Manh Hai LSTM Long Short Term Memory MOS Mean Opinion Score NLP Natural Language Processing RNN Recurrent Neural Network S2S Sequence-to-Sequence SER Speech Emotion Recognition SOTA State-Of-The-Art SPSS Statistical Parametric Speech Synthesis SUS Semantically Unpredictable Sentences TTH Tang Thanh Ha TTS Text To Speech VAE Variational Autoencoder WER Word Error Rate INTRODUCTION In recent years, TTS has become increasingly popular for general use, as it saves time and makes communication more accessible. One promising direction is the use of expressive speech synthesis, which aims to generate speech that conveys emotional nuances through prosody and other vocal cues. Expressive speech syn- thesis has the potential to revolutionize our interaction with technology by mak- ing it more natural and human-like.
It is a rapidly developing field, and there have been many recent advancements in this area worldwide. There are also a lot of technical difficulties related to expressive speech syn- thesis. One of the most significant issues is the requirement for huge amounts of training data. Generating speech that sounds human-like and conveys expres- siveness requires a large amount of data, and collecting and classifying this data can be time-consuming and expensive.
Another problem is the requirement for strong algorithms that can deal with variability in speech patterns. For example, expressive speech might differ depending on characteristics such as age, gender, and culture. Algorithms employed for expressive speech synthesis must be able to accommodate this variability and provide context-appropriate speech. It is cer- tainly possible to record large expressive datasets and apply the same complicated models, but the huge range of languages, speakers, and expressive and affect in- tensities makes this an ineffective experiment.
Besides that, modern expressive TTS models must make better use of the limited data they can train on and have integrated mechanisms that aid in generating expressive speech, as well as easy- to-interpret controls that are applicable in a variety of circumstances. In Vietnam, there has also been some progress in developing expressive speech synthesis, but it is still in the early stages of development. One of the main chal- lenges facing researchers in Vietnam is the lack of high-quality speech datasets, which can make it difficult to train accurate models for emotion speech synthesis. To the best of my knowledge, there has been little research on Vietnamese ex- pressive speech, such as [1]–[3]; however, none of the datasets contained in these papers have been made public.
Despite this, there is still room for improvement in this area. It is important to acquire a large, high-quality dataset for studying Vietnamese expressive speech. In this thesis, when referring to expressive speech, emotional speech is specifi- cally focused on. Emotional speech refers to the emotional state of the speaker and is conveyed through variations in tone, pitch, and volume.
Examples of emo- 1 tions conveyed through emotional speech include joy, sadness, anger, and fear. This type of speech helps communicate the speaker’s feelings and can also be used to elicit emotional responses from the listener. The main contributions of this thesis include: • Propose a semi-automatic pipeline to build Vietnamese emotional speech dataset. • Release two datasets of Vietnamese emotional speech using the described pipeline.
Analyze these Vietnamese emotional corpora and provide view- points. • Develop a model for emotional speech synthesis that is suitable for the spec- ified data objectives. The thesis is organized as follows: Chapter 1 provides an overview of speech features, speech synthesis, and expres- sive speech synthesis, focusing mostly on the technical side. Chapter 2 describes existing expressive datasets and some basic data processing steps.
It then presents TTH and LMH - two Vietnamese emotional speech datasets - and describes the strategy for the emotional corpus-building pipeline. Chapter 3 presents a baseline and proposed emotional speech synthesis model. Chapter 4 provides experimental results on various instances. Additionally, the effectiveness of the proposed corpus building pipeline is examined.
Part Conclusion concludes the thesis and outlines future works. THEORETICAL BACKGROUND This chapter provides an overview of speech features, speech synthesis, and ex- pressive speech synthesis, focusing mostly on the technical side.1 Non-emotional Features Speech is a signal that contains a lot of information. Depending on the purpose of the analysis, useful information will be extracted and analyzed. There are var- ious feature spaces that characterize speech data.
This section provides a simple overview of the features that are commonly used in Deep Learning architectures. a, Spectrogram A spectrogram is a visual representation of a signal’s frequency spectrum as it changes over time [4]. In the field of speech, spectrograms are often used to ana- lyze and change speech data. The spectrogram represents the frequency content of a spoken signal over time.
The x-axis indicates time, while the y-axis indicates frequency. At each location on the spectrogram, the color intensity corresponds to the amplitude or strength of the corresponding frequency component. Spectrograms are useful for judging how speech sounds, figuring out what pho- netic qualities they have, and learning about how speech sounds. They are also used in speech synthesis to change speech signals and make synthetic speech.
b, Mel Spectrogram Based on the idea that the human ear is more sensitive to some frequencies than others, this property tries to compress the representation of speech in the higher frequency domain. The mel scale is an experimental function that shows how sensitive the human ear is to different frequencies. The mel-spectrogram, which is based on the auditory-based mel-frequency scale, gives more frequency resolution than the spectrogram [5].1: An example of waveform, spectrogram, and mel-spectrogram. c, Acoustic/Prosodic Features Acoustic features are the physical qualities of the sound waves that the vocal tract makes [6].
These include parameters such as pitch, loudness, and duration of phonemes (units of sound). They provide information about the speaker’s emo- tions, attitudes, and intentions. Pitch describes the intonation of a sentence, while energy features cover the intensity of the uttered words. Duration stands for the speed of talking and the number of pauses.
Two more classes that do not directly belong to prosody are articulation (formants and bandwidths) and zero crossing rate. These deduced features are obtained by measuring statistical values of their corresponding extracted contours, such as mean, median, minimum, maximum, range, and variance. On the other hand, prosodic features are the patterns of stress, tone, and rhythm in speech [7]. These features play an important role in conveying meaning and emotion in human communication.
In speech synthesis, prosodic features need to be carefully modeled and synthesized in order to create a realistic-sounding synthetic speech.2 Emotional Features Valence and arousal are two key features used to describe emotions in speech [8].