1 of 14

JHU IWSLT 2023 Dialect Speech Translation System

2 of 14

Dialect Speech Translation

  • Dialect: A variety of a language spoken by a group of people, often in a specific geographic location

SodaPopCoke

[2]

aRmarm

[1]

[3]

Arabic Dialect Continuum

3 of 14

Dialect Speech Translation

  • Challenges:
    • Non-standard orthography
    • Rarely written
    • Low-resource
    • Conversational speech (disfluencies)

Arabic Dialect Continuum

4 of 14

Our IWSLT History

Focused on cascaded ST :

  • Finetuning seed MT/ASR models trained on mismatched MSA dialect in Cascaded
  • Synthetic Tunisian generated from MSA with bridged-back translation
  • Aggressively normalizing the source Tunisian transcripts for ASR and MT

2022

Focused on E2E and cascaded ST:

  • Dialectal transfer from large pre-trained models to improve translation in both E2E and Cascaded
  • E2E and Cascaded systems combination with minimum Bayes-Risk Decoding as post-processing
  • Reduce orthographic variations by pseudo-labeling source Tunisian
  • Further improve ASR by channel matching

2023

5 of 14

IWSLT 2023 Dialect Speech Translation Task

  • 2 conditions
    1. Constrained: 3-way parallel data:
      • 166h Tunisian Arabic
      • 212k utterances with manual Tunisian transcripts
      • 212k utterances with English translations
    2. Unconstrained (A.) + …
      • 1200h transcribed MSA Arabic broadcast news (MGB-2 corpus)
      • 250 h of telephone Levantine Arabic
      • Large pretrained MT models (mBart, NLLB-200)

6 of 14

ASR & MT

  • Conformer block Branchformer block

7 of 14

Experiments: Cascaded ST

  • ASR: Hybrid CTC/Attention Branchformer (Peng et al., 2022)
    • (A) train on constrained Tunisian Arabic
    • (B) pretrained on (MGB-2) + Levantine Arabic and finetuned on (A)
      • Mitigate orthographic variations: pseudo-labeling during finetuning
      • Match sampling rate: down sample the MGB-2 speech from 16kHz to 8kHz
      • Match the target channel: by additional telephone speech from the Levantine Arabic

8 of 14

Experiments: Cascaded ST

  • (A) Branchformer encoder, train on Tunisian-English data.
  • (B) incorporate large pretrained models: mBART and NLLB-200
    • mBART:
      • MSA, French, Italian, and Spanish, contribute loanwords to Tunisian (Zribi et al., 2014).
    • NLLB-200:
      • distilled 1.3B parameters,
      • Supports MSA and Tunisian Arabic, other closely related Maghrebi dialects (Moroccan, Maltese).

9 of 14

Experiments: Cascaded ST

  • (A1) Trained on Tunisian-English translations
  • (B2) mBART: finetuned on Tunisian-English translations
  • (B3) NLLB-200: finetuned on Tunisian-English translations

Best result

10 of 14

Experiments: End-to-End ST

  • End-to-end
    • Multitask learning approach: combines ASR and MT into differentiable E2E system

    • Hierarchical Audio Encoder: reorders

speech encoder

    • Incorporate pretrained large MT models

11 of 14

Experiments: System Combination

  •  

Utility function

likelihood

12 of 14

Resutls

Best overall result

E2E-ST system outperformed cascaded in constrained

Transcript normalization improves BLEU

cascaded outperformed E2E-ST in unconstrained

13 of 14

Conclusion

  • Presented dialectal transfer approaches from large pre-trained models to improve translation in both E2E and Cascaded ST settings
  • Pseudo-labeling and channel matching provided significant improvements for the ASR
  • Best results: system combination using Minimum Bayes-Risk decoding with COMET utility

14 of 14

Thank You�