1 of 11

2023.11.17

제네시스랩 AI연구팀

신재영

Arxiv 26 Sep 2023

2 of 11

Introduction

Introduction

  • With the growing user base, there's an increase in risks of casually created misinformation and malicious deepfakes.

​

  • Many deep learning-based detectors are vulnerable to spurious features and poor generalization to data from different domains.

​

  • We use HiFi-GAN as GANs play a key role in the adversarial game of detection and generation.

​

  • This paper proposes a collaborative training scheme where a generative model watermarks its output to make it easily detectable by a specific classifier

3 of 11

Background of watermarks

Introduction

si​ and sj​ are the spectral amplitudes at the i-th and j-th frequency bins, respectively

A and B are two sets of frequency bins

d is a parameter indicating the strength of the watermark

The process of embedding

a watermark into an audio signal

The process of detecting the watermark

4 of 11

Architecture

Overall Structure

5 of 11

Method

Loss Functions

​

    • HiFi-GAN losses

​

​

�

  • Denote generated speech signal as xgen = G(m)
    • G is the Generator model, and m is the mel-spectrum of the target waveform
  • D is the Discriminator model
  • E, is approximated by minibatch averages over batch elements and timesteps
  • M is a tensorized mel-filterbank matrix

6 of 11

Method

​

    • Watermark detector loss

​

�

  • Exactly the same objective as D
  • G can either share the WM objective or ignore it.
    • These two scenarios are called Collaborator and Observer, respectively

Loss Functions

7 of 11

Experiment

Datasets

​

    • Voice Cloning Toolkit (VCTK) corpus is used for all the speech data in the experiments

​

    • Noise data for augmentation consists of the noise subset from the MUSAN database

�

8 of 11

EER of experiment in testing conditions

Results

9 of 11

EER of experiment in testing conditions

Results

10 of 11

Contribution

Results

  • The results show that collaborative training consistently improves detection performance compared to baseline.

​

  • It would be useful to extend to full TTS voice cloning and know who generated the samples or what data was used to train the model.

​

  • It could be used as a benchmarking tool to evaluate how close our TTS output is to natural speech.

​

  • By analyzing the characteristics that make synthesized speech detectable, we might gain insights into aspects of the TTS output that need refinement, such as prosody or naturalness.

11 of 11

Q & A