Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Prosody Labeling with Phoneme-BERT and Speech F...

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest.
Avatar for hyama5 hyama5
August 24, 2025

Prosody Labeling with Phoneme-BERT and Speech Foundation Models

Avatar for hyama5

hyama5

August 24, 2025

More Decks by hyama5

Other Decks in Technology

Transcript

  1. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Speech Synthesis

    Workshop 2025,Aug. 24, 2025 Tomoki Koriyama (CyberAgent AI Lab)
  2. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Background: TTS

    w/ phoneme + prosody Text-to-speech Phoneme+prosody-to-speech “Is it sunny?” /IY Z IY T S AH H* N IY H%/ • Easy to collect data • • Mispronunciation & unintended prosody • Controllable phoneme & prosody Hard to collect annotated data Focus on automatic prosody annotation for prosody-controllable TTS 2
  3. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Background: Linguistic

    and speech foundation models • Linguistic foundation models improved prosody in TTS ◦ ◦ BERT [Xiao20, Kenter20] Phoneme-input BERT ▪ • Speech foundation model enables prosody prediction ◦ ◦ • PnG BERT [Wu21], phoneme-level (PL)-BERT [Li23] Whisper-encoder-based prosody prediction [Shirahata24] SSL-model-based label estimation [Kurihara24] Research question ◦ What if we combine linguistic and speech foundation models in automatic prosody annotation? 3
  4. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Purpose •

    Investigate automatic prosody annotation using both linguistic and speech foundation models • Approach ◦ Training on prosody-annotated speech data ▪ ◦ Linguistic foundation models ▪ ◦ Corpus of spontaneous Japanese (CSJ) PnG BERT, PL-BERT vs. one-hot Speech foundation models ▪ HuBERT, Whisper vs. melspectrogram, F0 4
  5. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Speech foundation

    models • Self-supervised learning (SSL)-based models ◦ Wav2vec2.0, HuBERT, WavLM, … [Hsu et al, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units”, 2021] • Whisper [Radford23] Speech foundation models are effective for extracting acoustic features 5
  6. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Phoneme-input language

    models • PnG BERT [Wu21] ◦ ◦ BERT with phoneme & grapheme input Predict masked tokens • PL-BERT [Li23] ◦ ◦ Predict graphemes as well as masked phonemes Not require grapheme input for inference Both PnG BERT and PL-BERT enhanced synthetic speech quality 6
  7. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Proposed model

    • Use hidden layer outputs of speech and linguistic foundation models • Input them to an annotation model to predict prosodic labels 7
  8. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Experimental conditions

    • Dataset ◦ The Corpus of Spontaneous Japanese (CSJ) ▪ ◦ Training data ▪ ◦ 33.2 h / 2.49 k utterances / 1.58 M phonemes Test data ▪ • Manually annotated prosody labels in the core data 1.3 h / 896 utterances Model ◦ ◦ Annotation model: 6-layer CNN, kernel size 5 Speech models ▪ ◦ HuBERT (base), Whsiper-encoder (small), melspec, F0, None Language models ▪ PnG BERT, PL-BERT, one-hot, None 8
  9. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prosody labels

    • ACC (accent labels) ◦ ◦ • HL ◦ • [, ]: low-high transition #, ? , %: low, high, high-low boundary pitch movement (BPM) • H, L: High-low PAU ◦ Y, N: Short-Pause existence BI (break index) ◦ ◦ 0, 1, 2, 3: no, word, accent-phrase, intonational-phrase boundary F, D: filler, disfluency 9
  10. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Results: Accent

    label (ACC) accuracy - HuBERT and Whsiper yielded higher accuracies than melspec and F0 PnG BERT and PL-BERT improved accuracy compared with one-hot 10
  11. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Results: Break

    index (BI) accuracy PnG BERT was better than PL-BERT on BI label prediction 11
  12. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Confusion matrix:

    Accent label (ACC) True Predicted - HuBERT reduced errors on boundary pitch movements (#, %, ?) 12
  13. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Confusion matrix:

    Break index (BI) True Predicted - PnG BERT enhanced accuracy on word boundaries (0 & 1) HuBERT reduced errors on intensities on phrase boundaries (2 & 3) 13
  14. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Prediction results

    Predicted labels are very close to the manually annotated ones 14
  15. Prosody Labeling with Phoneme-BERT and Speech Foundation Models TTS examples

    • Kono hana(鼻) wa chiisai (= This nose is small) TTS TTS w/ accent label /ko [ no # ha [ na wa # chi [ H sa ] i/ • Kono hana(花) wa chiisai (= This flower is small) TTS TTS w/ accent label /ko [ no # ha [ na ] wa # chi [ H sa ] i/ 15
  16. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Conclusions •

    Automatic prosody annotation model that uses the encoder hidden layers of both speech and linguistic foundation models • The combination of speech and linguistic models enhanced the prosody label prediction accuracies compared to using either acoustic or linguistic inputs alone • Future work ◦ ◦ Investigate the effect of predicted labels on TTS Apply the proposed model to other languages 16
  17. Prosody Labeling with Phoneme-BERT and Speech Foundation Models Training of

    PnG BERT & PL-BERT • Dataset ◦ • Tokenizer ◦ • cl-tohoku/bert-base-japanese-v3 Phonemes ◦ • CC-100 & Wikipedia corpus UniDic Training ◦ 1 M step optimization, batch size 64 18
  18. Prosody Labeling with Phoneme-BERT and Speech Foundation Models SSL-model comparison

    results Models trained on Japanese dataset seems better than English and multilingual ones 19