INESCO Dataset: indonesian expressive speech corpus for emotion speech synthesis.
Source: PubMed, NCBI / U.S. National Library of Medicine
A speech corpus is a crucial component in the development of language models in natural language processing (NLP). A high-quality corpus is a crucial factor in the quality of NLP study, including text-to-speech (TTS), also known as speech synthesis. In Indonesia, developing a high-quality corpus is a priority to ensure that speech synthesis technology better supports digital communications, the creative industries, and public services. Indonesian is still categorized as a low-resourced language, which means its use in digital communication remains limited. Therefore, this study addresses this gap by developing an Indonesian Expressive Speech Corpus (INESCO). INESCO was developed by selecting sentences from various literary sources, including textbooks, magazines, novels, newspapers, movies, and websites. The selected sentences consist of commonly used expressions in everyday conversation. The expressive Indonesian speech corpus contains three emotional styles: happiness (211 sentences), anger (190), and sadness (199 sentences). The total number of sentences in the Indonesian expressive speech corpus were 600. The Indonesian expressive speech corpus was recorded by four professional speakers, consisting of two females and two males. The selected voice artists were professional theater artists who were members of the Indonesian Arts Council, namely Teater Api from Surabaya, East Java, Indonesia and Teater Payung Hitam from Bandung, West Java, Indonesia. The voice artists were s
Abstract
A speech corpus is a crucial component in the development of language models in natural language processing (NLP). A high-quality corpus is a crucial factor in the quality of NLP study, including text-to-speech (TTS), also known as speech synthesis. In Indonesia, developing a high-quality corpus is a priority to ensure that speech synthesis technology better supports digital communications, the creative industries, and public services. Indonesian is still categorized as a low-resourced language, which means its use in digital communication remains limited. Therefore, this study addresses this gap by developing an Indonesian Expressive Speech Corpus (INESCO). INESCO was developed by selecting sentences from various literary sources, including textbooks, magazines, novels, newspapers, movies, and websites. The selected sentences consist of commonly used expressions in everyday conversation. The expressive Indonesian speech corpus contains three emotional styles: happiness (211 sentences), anger (190), and sadness (199 sentences). The total number of sentences in the Indonesian expressive speech corpus were 600. The Indonesian expressive speech corpus was recorded by four professional speakers, consisting of two females and two males. The selected voice artists were professional theater artists who were members of the Indonesian Arts Council, namely Teater Api from Surabaya, East Java, Indonesia and Teater Payung Hitam from Bandung, West Java, Indonesia. The voice artists were selected to record the corpus. Based on the criteria of being allowed to have a regional accent and having more than five years of experience, they were recorded in a soundproof studio with separate recordings and operator rooms to minimize background noise. The recording process used a condenser microphone (Neumann U87 and Studio Project B1) that was equipped with a pop filter. The speakers were asked to stand 2-3 cm in front of the microphone. The microphone was connected to Presonus Studio One version 4.1.4 software for recording and editing. Recorded speech corpus was under configuration with sampling frequency of 44.1 kHz, channel input/output mono, 16 bits/sample and using format ".wav" file. During the recording process, the chair of the Indonesian Arts Council supervised to ensure that every sentence recorded in the database conveyed the correct meaning and evaluated whether all the speakers' balance was maintained across the three types of emotions. The phonetic annotation process was performed using the Wavesurfer speech software. The emotional validation method was implemented through a set of preliminary experiments on the end-to-end TTS system for Indonesian to demonstrate the dataset's usability. Only files with a mean opinion score (MOS) score of 3.5 or higher were used in the INESCO dataset. The final INESCO dataset compiled 2400 audio files with a text transcription. The total size of the files was 1.33 GB. The total duration of this speech corpus was 3 h, 1 min, and 49 s. This dataset contains comprehensive, carefully curated data to enable reproducible experimentation in end-to-end TTS systems and to support research in speech synthesis, speech recognition, and other areas of natural language processing. The INESCO dataset is publicly accessible and useful for research, academic, and educational purposes.
