1 of 19

X International conference

“Information Technology and Implementation” (IT&I-2023)

Kyiv, Ukraine

1

Using Parallelized Neural Networks to Detect Falsified Audio

Information in Socially Oriented Systems

​

Artem Khovrat, Volodymyr Kobziev and Sergiy Yakovlev

​

Dedicated to the tenth anniversary of the Faculty of Information Technology

2 of 19

Introduction

In recent decades, technologies capable of falsifying information have reached the level where the need to detect forgeries in socially oriented systems is being discussed at the legislative level. In particular, in the case of video information, the distortion has not yet reached the required level, but audio falsification has recently been able to go beyond simple identification, that is, using human hearing.

​

How does this threat manifest itself?

​

Use of synthetically generated audio can have direct impact on military operations, including voice spoofing to create a trap and spoofing radio signals.

​

What to do in this case?

​

Find a method that allows us to effectively recognize the fact of forgery of an audio message.

​

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

3 of 19

Object

Neural networks with elements of convolutional and recurrent architectures.

​

Purpose

Development of an effective model for determining the fact of falsification of audio data, using MapReduce technology.

​

Tasks

    • determine the features of audio in socially oriented systems;
    • analyse the international experience of determining the fact of falsification for various types of information;
    • based on the conducted analysis, form a set of limitations and assumptions, and define models that will be used in further research;
    • carry out a description of algorithms that would allow the reprocessing of audio information both in the form converted to text and in the form of a signal;
    • review selected architectures of neural networks and determine their main hyperparameters;
    • form an experiment plan and a multi-criteria selection task that would allow determining the expediency of using MapReduce and choosing the most effective classification algorithm;
    • analyse the results of the experiment and form appropriate conclusions.

​

​

​

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

About Study

4 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

Similar Investigations

1

BREUR, EILAT,

WEINSBERG

Friend or Faux: Graph-Based Early Detection of Fake Accounts on Social Networks (2020)

​

ALONSO, VILARES,

GOMEZ, VILARES

Sentiment Analysis for Fake News Detection (2021)

NAJAR, ZAMZAMI,

BOUGULIA

​

Fake News Detection Using Bayesian Inference (2019)

CHOUDHARY,

ARORA

Linguistic feature based learning model for fake news

detection and classification (2021)

TOLOSANA, VERA,

FIERREZ, MORALES,

ORTEGA-GARCIA

Deepfakes and beyond: A Survey of face manipulation and fake detection (2020)

​

5 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

Domain Analysis

1

Ways of falsification

    • synthetic creation of audio with the help of artificial intelligence;
    • composition of existing sound tracks to distort the essence of the original information;
    • contextual distortion.

​

Features of falsified audio

​

​

​

1

2

4

5

3

LARGE NUMBER

of rhetorical questions

with a condemnatory meaning

ABSCENCE

of negative constructions to reduce cognitive load

ORALS LIMITATION I

some features may not manifest themselves

ORALS LIMITATION II

incorrect audio recognition or pronunciation features

USING APPEALS

and incitements in inappropriate contexts

6 of 19

10

10

1

2

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

Audio as Text

1

It is worth highlighting the following features typical for the subject area:

CONVERT AUDIO

TO TEXT

1

2

CONVERSION

pandas

3

TEXT TOKENIZATION

4

TEXT

CLEANING

5

TEXT

STEMMING

6

text Lematization

7

CLEANING

words

13

8

12

9

Algorithm

11

10

BM25

CHARACTERISTICS

POLARITY

words

EXTERNAL CHARACTERISTICS

​

MESSAGE

WEIGHTS

DEGREE OF RELIABILITY

AGGREGATION OF RESULT

SPEECH TO TEXT:

    • quality of recordings for training;
    • lack of data for model formation;
    • ignoring pronunciation defects;
    • correct processing of dialectics, neologisms, abbreviations.

SPECIAL ISUESS:

    • abbreviations;
    • pauses between word;
    • quality of recordings;
    • neologisms and dialectisms.

​

7 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

Audio as Signal

1

The Mahalanobis distance can be determined as follows using the formula:

1

SPLIT INTO

200 MS WINDOW

2

FIND DISTANCE BY MAHALANOBIS

3

IDENTIFY

VOCALIZED AREAS

4

AUGMENT DATA

BY VAR

5

CONVERT INTO

NUMERIC FORM

0

GENERATE

DATASET

Based on the Gaussian distribution formula for a 200 ms window, we obtained the following set of rules:

8 of 19

CLASSIC

LSTM

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

LSTM Architecture

1

BIDIRECTIONAL LSTM

In the classical architecture of recurrent neural networks, the problems of vanishing and exploding gradients may arise. To solve them, the LSTM architecture was created.

9 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

CNN Architecture

1

1

2

4

5

3

DEPTH

equal to the number of filters used in the convolution (5)

STRIDE

determines the size of the step with which the passage is made (1)

BIAS

adjusts the shift in the convolution

operation in the corresponding layer (None)

​

KERNEL SIZE

determines the size of the filters used to analyze the text (via cross-validation determined 4)

PADDING

controls the addition of non-signifi-cant zeros if necessary (None)

10 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

RCNN Architecture

1

The last architecture to be considered will be the combination of several convolutional networks with bidirectional LSTM networks

11 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

MapReduce Technology

1

Based on the specifics of the subject area and testing the effectiveness, it was decided to use the Hadoop version of MapReduce:

12 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

Experimental Environment

1

In general, the experimental environment has several key characteristics, apart from the data sets themselves: plan of the experiment; multi-criteria selection.

​

ACCURACY

TIME-SAVING

VOLUME-SAVING

Determined by the combination of F1-score and Precision, normalized between 0 and 1

Importance: 16

Of model training for same capacities for two type of audio analyze process

Importance: 2

Minimum permissible amount of data volume to achieve Accuracy equal to 90%

Importance: 2

CONTEXT

The possibility of taking into account the context determined by special rules

Importance: 10

13 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

1

Models Implementation: Modification

Algorithm uses for a large number of different modules and Python libraries:

1

PACKAGES:

    • re;
    • polars;
    • scikit-learn;
    • numpy.

2

NLTK MODULES:

    • punkt;
    • stopwords;
    • wordnet;
    • vader_lexicon.

​

14 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

1

Models Implementation: MapReduce

Mapper (VAR)

Reducer (CNN)

15 of 19

Model

Time-Saving

Accuracy

Volume-Saving

CNN

1.00

0.90

0.29

RNN

0.90

0.90

0.00

LSTM

0.59

0.93

0.53

BiLSTM

0.31

0.96

0.82

RCNN

0.00

0.97

1.00

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

1

Experiment Results I

Criteria values for each alternative for audio as signal

(min value - 0, max value - 1):

16 of 19

Model

Time-Saving

Accuracy

Volume-Saving

Context

CNN

1.00

0.91

0.29

0.60

RNN

0.91

0.91

0.00

0.40

LSTM

0.53

0.93

0.53

0.80

BiLSTM

0.27

0.96

0.82

1.00

RCNN

0.00

0.96

1.00

1.00

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

1

Experiment Results II

Criteria values for each alternative for audio as text

(min value - 0, max value - 1):

17 of 19

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

1

Experiment Results III

After processing the above results and calculating the values of linear additive convolution with weighting coefficients, the following was established:

AUDIO AS SIGNAL

MAP REDUCE

The most effective model is BiLSTM, however, when parallelized, a similar efficiency result is obtained by RCNN.

In the case of BiLSTM, the acceleration is 2, in the case of RCNN it is about 3.5 and 4.3 for 3 and 4 nodes, respectively

AUDIO AS TEXT

As in the another case The most effective model is BiLSTM (when parallelized, a similar efficiency result is obtained by RCNN).

18 of 19

Conclusion

The current work aimed to develop an effective model for determining the fact of falsification of

audio data, using MapReduce technology.

​

Data

Audio as text on the election process in Ukraine in 2019 and the full-scale invasion of the Russian

Federation on the territory of Ukraine and the similar data as signal.

​

Models

Based on the review of the analyzed studies, it was decided to focus on: classical CNN, classical RNN, LSTM, BiLSTM, hybrid neural network combining several convolutional networks with a bidirectional

recurrent network with long-term memory.

​

Result

In the course of the experiments, it was found that the BiLSTM is the most effective, although it loses in speed to less complex models. It is found that the gain in reprocessing time saving when using MapReduce technology can reach 4.3 in the case of text and 4 in the case of signal.

​

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine

19 of 19

References

JAIN

Classifying Fake News

PAI

CNN vs. RNN vs. ANN

PENNINGTON

​

Global Vectors for Word Representation

RISDAL

Getting Real about Fake News

MCUBA

ET AL.

Deep Learning Methods on

Deepfake Audio Detection

​

BATALLIER ET AL.

Signal Detection Approach to Under-standing the Identification of Fake News

​

​

AFANASIEVA

ET AL.

​

Application of Neural Networks to

Identify of Fake News

​

MTASHER

ET AL.

Generate Poems and Letters Using an Iterative Neural Network

​

BANASL

ET AL.

Real-Time Advanced Computational Intelligence for Deep Fake Video Detection

​

AMIDI

Recurrent Neural Networks cheatsheet

DB CAMP

​

Long Short-Term Memory Networks (LSTM)

​

DOMAIN ANALYSIS

TECHNICAL IMPLEMENTATION

REDDY

Fake News Detection

Some of the additional source to create the presentation and study:

Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine