X International conference
“Information Technology and Implementation” (IT&I-2023)
Kyiv, Ukraine
1
Using Parallelized Neural Networks to Detect Falsified Audio
Information in Socially Oriented Systems
Artem Khovrat, Volodymyr Kobziev and Sergiy Yakovlev
Dedicated to the tenth anniversary of the Faculty of Information Technology
Introduction
In recent decades, technologies capable of falsifying information have reached the level where the need to detect forgeries in socially oriented systems is being discussed at the legislative level. In particular, in the case of video information, the distortion has not yet reached the required level, but audio falsification has recently been able to go beyond simple identification, that is, using human hearing.
How does this threat manifest itself?
Use of synthetically generated audio can have direct impact on military operations, including voice spoofing to create a trap and spoofing radio signals.
What to do in this case?
Find a method that allows us to effectively recognize the fact of forgery of an audio message.
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Object
Neural networks with elements of convolutional and recurrent architectures.
Purpose
Development of an effective model for determining the fact of falsification of audio data, using MapReduce technology.
Tasks
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
About Study
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Similar Investigations
1
BREUR, EILAT,
WEINSBERG
Friend or Faux: Graph-Based Early Detection of Fake Accounts on Social Networks (2020)
ALONSO, VILARES,
GOMEZ, VILARES
Sentiment Analysis for Fake News Detection (2021)
NAJAR, ZAMZAMI,
BOUGULIA
Fake News Detection Using Bayesian Inference (2019)
CHOUDHARY,
ARORA
Linguistic feature based learning model for fake news
detection and classification (2021)
TOLOSANA, VERA,
FIERREZ, MORALES,
ORTEGA-GARCIA
Deepfakes and beyond: A Survey of face manipulation and fake detection (2020)
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Domain Analysis
1
Ways of falsification
Features of falsified audio
1
2
4
5
3
LARGE NUMBER
of rhetorical questions
with a condemnatory meaning
ABSCENCE
of negative constructions to reduce cognitive load
ORALS LIMITATION I
some features may not manifest themselves
ORALS LIMITATION II
incorrect audio recognition or pronunciation features
USING APPEALS
and incitements in inappropriate contexts
10
10
1
2
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Audio as Text
1
It is worth highlighting the following features typical for the subject area:
CONVERT AUDIO
TO TEXT
1
2
CONVERSION
pandas
3
TEXT TOKENIZATION
4
TEXT
CLEANING
5
TEXT
STEMMING
6
text Lematization
7
CLEANING
words
13
8
12
9
Algorithm
11
10
BM25
CHARACTERISTICS
POLARITY
words
EXTERNAL CHARACTERISTICS
MESSAGE
WEIGHTS
DEGREE OF RELIABILITY
AGGREGATION OF RESULT
SPEECH TO TEXT:
SPECIAL ISUESS:
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Audio as Signal
1
The Mahalanobis distance can be determined as follows using the formula:
1
SPLIT INTO
200 MS WINDOW
2
FIND DISTANCE BY MAHALANOBIS
3
IDENTIFY
VOCALIZED AREAS
4
AUGMENT DATA
BY VAR
5
CONVERT INTO
NUMERIC FORM
0
GENERATE
DATASET
Based on the Gaussian distribution formula for a 200 ms window, we obtained the following set of rules:
CLASSIC
LSTM
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
LSTM Architecture
1
BIDIRECTIONAL LSTM
In the classical architecture of recurrent neural networks, the problems of vanishing and exploding gradients may arise. To solve them, the LSTM architecture was created.
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
CNN Architecture
1
1
2
4
5
3
DEPTH
equal to the number of filters used in the convolution (5)
STRIDE
determines the size of the step with which the passage is made (1)
BIAS
adjusts the shift in the convolution
operation in the corresponding layer (None)
KERNEL SIZE
determines the size of the filters used to analyze the text (via cross-validation determined 4)
PADDING
controls the addition of non-signifi-cant zeros if necessary (None)
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
RCNN Architecture
1
The last architecture to be considered will be the combination of several convolutional networks with bidirectional LSTM networks
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
MapReduce Technology
1
Based on the specifics of the subject area and testing the effectiveness, it was decided to use the Hadoop version of MapReduce:
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
Experimental Environment
1
In general, the experimental environment has several key characteristics, apart from the data sets themselves: plan of the experiment; multi-criteria selection.
ACCURACY
TIME-SAVING
VOLUME-SAVING
Determined by the combination of F1-score and Precision, normalized between 0 and 1
Importance: 16
Of model training for same capacities for two type of audio analyze process
Importance: 2
Minimum permissible amount of data volume to achieve Accuracy equal to 90%
Importance: 2
CONTEXT
The possibility of taking into account the context determined by special rules
Importance: 10
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
1
Models Implementation: Modification
Algorithm uses for a large number of different modules and Python libraries:
1
PACKAGES:
2
NLTK MODULES:
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
1
Models Implementation: MapReduce
Mapper (VAR)
Reducer (CNN)
Model | Time-Saving | Accuracy | Volume-Saving |
CNN | 1.00 | 0.90 | 0.29 |
RNN | 0.90 | 0.90 | 0.00 |
LSTM | 0.59 | 0.93 | 0.53 |
BiLSTM | 0.31 | 0.96 | 0.82 |
RCNN | 0.00 | 0.97 | 1.00 |
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
1
Experiment Results I
Criteria values for each alternative for audio as signal
(min value - 0, max value - 1):
Model | Time-Saving | Accuracy | Volume-Saving | Context |
CNN | 1.00 | 0.91 | 0.29 | 0.60 |
RNN | 0.91 | 0.91 | 0.00 | 0.40 |
LSTM | 0.53 | 0.93 | 0.53 | 0.80 |
BiLSTM | 0.27 | 0.96 | 0.82 | 1.00 |
RCNN | 0.00 | 0.96 | 1.00 | 1.00 |
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
1
Experiment Results II
Criteria values for each alternative for audio as text
(min value - 0, max value - 1):
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
1
Experiment Results III
After processing the above results and calculating the values of linear additive convolution with weighting coefficients, the following was established:
AUDIO AS SIGNAL
MAP REDUCE
The most effective model is BiLSTM, however, when parallelized, a similar efficiency result is obtained by RCNN.
In the case of BiLSTM, the acceleration is 2, in the case of RCNN it is about 3.5 and 4.3 for 3 and 4 nodes, respectively
AUDIO AS TEXT
As in the another case The most effective model is BiLSTM (when parallelized, a similar efficiency result is obtained by RCNN).
Conclusion
The current work aimed to develop an effective model for determining the fact of falsification of
audio data, using MapReduce technology.
Data
Audio as text on the election process in Ukraine in 2019 and the full-scale invasion of the Russian
Federation on the territory of Ukraine and the similar data as signal.
Models
Based on the review of the analyzed studies, it was decided to focus on: classical CNN, classical RNN, LSTM, BiLSTM, hybrid neural network combining several convolutional networks with a bidirectional
recurrent network with long-term memory.
Result
In the course of the experiments, it was found that the BiLSTM is the most effective, although it loses in speed to less complex models. It is found that the gain in reprocessing time saving when using MapReduce technology can reach 4.3 in the case of text and 4 in the case of signal.
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine
References
JAIN
Classifying Fake News
PAI
CNN vs. RNN vs. ANN
PENNINGTON
Global Vectors for Word Representation
RISDAL
Getting Real about Fake News
MCUBA
ET AL.
Deep Learning Methods on
Deepfake Audio Detection
BATALLIER ET AL.
Signal Detection Approach to Under-standing the Identification of Fake News
AFANASIEVA
ET AL.
Application of Neural Networks to
Identify of Fake News
MTASHER
ET AL.
Generate Poems and Letters Using an Iterative Neural Network
BANASL
ET AL.
Real-Time Advanced Computational Intelligence for Deep Fake Video Detection
AMIDI
Recurrent Neural Networks cheatsheet
DB CAMP
Long Short-Term Memory Networks (LSTM)
DOMAIN ANALYSIS
TECHNICAL IMPLEMENTATION
REDDY
Fake News Detection
Some of the additional source to create the presentation and study:
Information Technology and Implementation, November 20, 2023, Taras Shevchenko National University of Kyiv, Kyiv, Ukraine