w/ phoneme + prosody Text-to-speech Phoneme+prosody-to-speech “Is it sunny?” /IY Z IY T S AH H* N IY H%/ • Easy to collect data • • Mispronunciation & unintended prosody • Controllable phoneme & prosody Hard to collect annotated data Focus on automatic prosody annotation for prosody-controllable TTS 2
and speech foundation models • Linguistic foundation models improved prosody in TTS ◦ ◦ BERT [Xiao20, Kenter20] Phoneme-input BERT ▪ • Speech foundation model enables prosody prediction ◦ ◦ • PnG BERT [Wu21], phoneme-level (PL)-BERT [Li23] Whisper-encoder-based prosody prediction [Shirahata24] SSL-model-based label estimation [Kurihara24] Research question ◦ What if we combine linguistic and speech foundation models in automatic prosody annotation? 3
Investigate automatic prosody annotation using both linguistic and speech foundation models • Approach ◦ Training on prosody-annotated speech data ▪ ◦ Linguistic foundation models ▪ ◦ Corpus of spontaneous Japanese (CSJ) PnG BERT, PL-BERT vs. one-hot Speech foundation models ▪ HuBERT, Whisper vs. melspectrogram, F0 4
• Dataset ◦ The Corpus of Spontaneous Japanese (CSJ) ▪ ◦ Training data ▪ ◦ 33.2 h / 2.49 k utterances / 1.58 M phonemes Test data ▪ • Manually annotated prosody labels in the core data 1.3 h / 896 utterances Model ◦ ◦ Annotation model: 6-layer CNN, kernel size 5 Speech models ▪ ◦ HuBERT (base), Whsiper-encoder (small), melspec, F0, None Language models ▪ PnG BERT, PL-BERT, one-hot, None 8
label (ACC) accuracy - HuBERT and Whsiper yielded higher accuracies than melspec and F0 PnG BERT and PL-BERT improved accuracy compared with one-hot 10
• Kono hana(鼻) wa chiisai (= This nose is small) TTS TTS w/ accent label /ko [ no # ha [ na wa # chi [ H sa ] i/ • Kono hana(花) wa chiisai (= This flower is small) TTS TTS w/ accent label /ko [ no # ha [ na ] wa # chi [ H sa ] i/ 15
Automatic prosody annotation model that uses the encoder hidden layers of both speech and linguistic foundation models • The combination of speech and linguistic models enhanced the prosody label prediction accuracies compared to using either acoustic or linguistic inputs alone • Future work ◦ ◦ Investigate the effect of predicted labels on TTS Apply the proposed model to other languages 16