1 of 81

Multimodal Approach for Novelty Aware Emotion Recognition with Situational Knowledge

Presented By

Mijanur R Palash

PhD Candidate

Committee:

Dr. Bharat Bhargava (Advisor)

Dr. Chunyi Peng

Dr. Jianguo Wang

Dr. Vaneet Aggarwal

1

July 20, 2023

2 of 81

Agenda

2

Background and Motivation

Contribution 1: SAFER (Emotion Recognition From Face)

Contribution 2: EMERSK (Multimodal Emotion Recognition)

Contribution 3: CoNERS (Novelty Aware Emotion Recognition)

Question & Answer

3 of 81

Background and Motivation

3

How can we help??

The USA: in a mental health crisis

Mental health: influences gun violence, school shooting, suicide etc.

4 of 81

Background and Motivation

4

  • Close relation between emotion and mental health
  • Changes in emotions over time used for:
    • Trigger identification
    • Early sign of instability
    • Preventive steps

  • Our idea of help:
    • Automated emotion recognition which can be used for:
      • Automated monitoring
      • Advance warning
      • Alarm triggering

5 of 81

Emotion Indicators

5

Visual

    • Facial expression
    • Posture
    • Gait

Non-visual

    • Speech
    • Text
    • Brain scan

Emotions can be conveyed through both visual and non-visual indicators.

6 of 81

Challenges in Emotion Recognition

6

    • Providing high accuracy in emotion recognition
    • Most focused area

Accuracy

    • Giving transparent explanations of the results
    • Lack of focus

Explainability

    • Detecting and adapting to novel situations
    • Lack of focus

Novelty Handling

7 of 81

Existing Works

7

    • Heavily focused on facial emotion recognition (FER)

Face Based

    • Use only one or two modes

Unimodal/Bimodal

    • Not focused on explainable output

Not Explainable

    • Not designed to handle novelty

Novelty

8 of 81

Contributions

8

SAFER: Improved facial emotion recognition

EMERSK: Explainable multimodal emotion recognition

CoNERS: Novelty detection and handling

9 of 81

9

SAFER: Situation Aware Facial Emotion Recognition

10 of 81

Problem Statement

10

Can we improve the facial emotion recognition?

Bias in Amazon AI gender classification

Facial expression of emotions

  • Face: important medium of emotion
  • Subject to bias: need generalization

11 of 81

SAFER Architecture

11

11

Face feature extraction

Background feature extraction

Place feature extraction

Classification network

12 of 81

Face Feature Extraction: Face Detection

12

BlazeFace [11] for face detection

    • Identifies key points
    • Generates face mesh

Face detection

13 of 81

Face Feature Extraction: Feature Types

13

Face Feature Types

Action unit (AU) features

Visible features

Deep features

Face feature extraction module

14 of 81

Face Feature Extraction: Action Unit (AU) Features

14

AU ID

AU Name

Points

1

Inner brow raiser

Above inner brow

6

Cheek raiser

At cheek center

24

Lip pressor

Bottom lip center

Action units for “Sadness”

Action unit features generation

  • AUs: set of face muscles that corresponds to specific expressions
  • BlazePose: computer vision model used to detect the centers of the AUs

15 of 81

Face Feature Extraction: Visible Features

15

Feature type

Description

Width

Left eye

Distance

Left and right eyes

Angle

Left eye with right eye and mouth

Visible features

    • Reflect physical changes of face parts with emotion
    • Measure as width, distance and angle

Visible features

16 of 81

Face Feature Extraction: Deep Features

16

16

  • Deep features: representations from the deeper layers of a CNN
  • Transfer learning:
  • Knowledge gained in one task applied to improve the performance of a related but different task
  • Resnet-50 pre-trained on ImageNet dataset (14 million samples)

17 of 81

Background Feature Extraction

17

17

  • Background: source of important contextual information
  • Process:
    • subject removal
    • convolutional feature extractor

18 of 81

Place Feature Extraction

18

18

  • Places are associated with emotion:
    • garden: happiness, cemetery: sadness
  • Provides additional information in emotion recognition and explanation generation
  • Pre-trained Model
    • AlexNet
  • Place dataset [23]
    • 10 million labeled images
    • 205 place categories

Place category: “Bedroom”

19 of 81

Evaluation Setup

19

19

 

20 of 81

Evaluation Setup: Related Works

20

20

Name

Method

Limitations

Wen et al. [34]

Ensemble CNN

Low accuracy

Dhankar et al. [17]

ResNet-50

Low accuracy

Renda et al. [35]

Ensemble CNN

Low accuracy

Gan et al. [16]

Soft

Label boosting+ ECNN

Not emphasized on all face points

A-C [18]

Adaptive correlation-based loss

Orthogonal work

Lee et al. [6]

Two stream architecture with adaptive fusion

Not well generalized as mainly focused on CAER-S dataset

Kosti et al. [5]

Dual stream CNN

Not well generalized as mainly focused on EMOTIC dataset

Li et al. [33]

Relational region-level analysis with Body-Object and Body-Part attention+ GCN

Accuracy can be improved

21 of 81

Evaluation Setup: Dataset

21

21

FER-2013: 3.2K posed images

CK+: 593 posed and spontaneous videos

AffectNets: 450K spontaneous image

CAER-S: 70K image from TV shows

RAF-DB: 30K diverse face images

FABO: 206 posed videos

Sample images from the datasets

22 of 81

Experimental Results and Findings

22

22

Research question: Does safer improve accuracy?

22

  • X axis: Name of the method ; Y axis: Accuracy reported by them in the dataset
  • The higher the bar, the better!

Results on FER-2013

Results on CAER-S

Findings: SAFER improves accuracy and outperforms state-of-the-art methods.

23 of 81

Experimental Results and Findings

23

23

Research question: Does Safer Generalizes Result?

23

23

Findings: SAFER shows high accuracy in all six datasets which proves good generalization.

Results on various FER datasets

24 of 81

Experimental Results and Findings

24

Research question: Which emotions are easy, and which are difficult to identify

Findings: Unbalanced classes and shared facial expressions degrade performance.

    • Lower sample number, underfitting
    • ‘Happiness’: High accuracy, ‘Disgust : Low accuracy
    • ‘Happiness’: 7000 samples, ‘Disgust’: 436 samples

Effect of sample numbers

    • Share some of the facial expressions
    • ‘Disgust’ and ‘Anger’, ‘Sad’ and ‘Neutral’
    • Misclassified for each other

Effect of related classes

Confusion Matrix on FER-2013

25 of 81

Ablation Study

25

25

Research question: Which combination offers best results

25

25

Experiments on CAER-S dataset

Findings: Best result is achieved when F,B and P are combined.

F: Face

B: Background

P: Place

26 of 81

Issues with Current Datasets

26

26

26

  • Bias:
    • gender and racial
  • Quality concern
  • No face mask

Class

Sample with male subject

Anger

70%

Happy

60%

Gender bias in FER-2013

Example of bad samples

27 of 81

Proposed Dataset

27

27

Research question: How can we improve the FER training?

27

A new dataset-

    • Seven emotion classes
      • Balanced
      • 3000+ sample each
    • Gender and ethnically diverse
    • Section for masked sample

28 of 81

Research Contribution of SAFER

28

A novel face feature extraction module

A novel facial emotion recognition system with background and place features

A detailed evaluation framework to prove the high accuracy and generalizability

A novel dataset for FER with masked subjects

29 of 81

Contributions

29

SAFER: Improved facial emotion recognition

EMERSK: Explainable multimodal emotion recognition

CoNERS: Novelty detection and handling

30 of 81

Problem Statement

30

Issues with the facial expression

    • Face covering
    • Intentional misleading

Use of multiple modalities can help!!

31 of 81

EMERSK

31

EMERSK

Explainable

Multimodal

Situational Knowledge

32 of 81

EMERSK: Architecture

32

Facial Module

Posture Module

Gait Module

Background Modul

Explanation Module

33 of 81

Face Module

33

Two-stream architecture

CNN

Attention based encoder-decoder network

34 of 81

Posture Module

34

Body modeling and posture detection

    • Kinematic representation of human body
      • Collection of joints
    • BlazePose for body point detection

Posture example

Body points

35 of 81

Posture Module

35

    • Calculated from the body points in the form of distance, angle, area etc.

Visible feature

generation

    • Deep representation using convolutional network

Deep feature generation

Visible feature Type

Description

Angle

At neck by both shoulders

Distance

Between right hand and hips joint

Area

Triangle between both hands and neck

36 of 81

Gait Module

36

Two-stream architecture

Upper stream: LSTM

Lower stream: 3D CNN

37 of 81

Background Module

37

Subject removal

Deep feature extraction

38 of 81

Place Module

38

38

Place dataset

Pre-trained AlexNet

Place category: “Bedroom”

39 of 81

Adjective-Noun Pair (ANP) Module

39

39

SentiBank 2.0: CNN based ANP classifier

Trained on one million images from Flickr

Crying baby

Colorful butterfly

40 of 81

Evaluation Setup

40

40

 

41 of 81

Evaluation Setup: Similar Works

41

41

Name

Method

Limitations

Explainable?

CMEFA [52]

Broad deep learning fusion network (BDFN) on Face and Posture

Limited evaluation

No

Bhatia et al. [57]

Layered LSTM on gait

Gait mode only

No

Kosti et al. [5]

Dual stream CNN on body and background

Considers whole body as a single mode

No

Lee et al. [6]

Two stream CNN with adaptive fusion on face and background

Posture and gait not considered

No

Santosh et al. [63]

ConvLSTM

Not modular, treats the video as a single mode

No

Tahghighi et al. [47]

HOG-KLT+ SVM

Considers whole body as a single mode

No

42 of 81

Experimental Results and Findings: Face Module

42

42

Research question: Can face module perform standalone?

42

Findings: Face module provide superior standalone performance.

Results on FER-2013

Results on CAER-S

43 of 81

Experimental Results and Findings: Posture Module

43

43

Research question: Can posture module perform standalone?

43

Findings: Posture module provide comparable and generalized standalone performance.

Results on FABO dataset

Results on various datasets

44 of 81

Experimental Results and Findings: Gait Module

44

44

Research question: Can gait module perform standalone?

44

Findings: Gait module provide superior standalone performance.

Results on FABO dataset

45 of 81

Experimental Results and Findings: Multimodal Operation

45

45

Research question: Does multimodal improve performance?

45

Findings: Multimodal provides superior performance than state-of-the-arts.

Experiments on GroupWalk dataset

Experiments on GEMEP dataset

46 of 81

Ablation Study

46

46

Research question: What is the best combination of the modes?

46

Findings: Face is the most expressive mode and multimodal beats standalone methods.

Experiments on GroupWalk dataset

47 of 81

Computational Cost

47

47

Research question: What is the computational cost of going multimodal?

47

Findings: Face is the fastest mode and gait is the slowest.

Results on EWALK dataset

48 of 81

Explanation Generation

48

Research question: How do we explain the output?

48

Individual mode result

Place type

Adjective-Noun pair

Average emotion

Explanation: “Emotion output is “happiness”. The place is “nursery_classroom”, it is “positive” environment, with “creative_work” and “smiling_kid”. Subject face is: “happy”,and posture is: “happy””.

49 of 81

Research Contribution of EMERSK

49

49

A modular architecture for emotion recognition from multiple modes

A novel approach for situational

knowledge generation

A novel approach for explanation generation

50 of 81

Contributions

50

SAFER: Improved facial emotion recognition

EMERSK: Explainable multimodal emotion recognition

CoNERS: Novelty detection and handling

51 of 81

What About Novelty?

51

51

Research question: What happens with frequent novel or unexpected samples?

51

Need to detect and handle novelty

Novelty is defined as a new or unusual instance that deviate from the expected norm!

Example of Novelty: A rhino freely roaming the streets of west Lafayette

52 of 81

CoNERS: Continuous Learning Based Novelty Aware Emotion Recognition System

52

Continuous Learning Loop

Emotion Recognition

Novelty Detection

Novelty Handling

Retraining

52

53 of 81

CoNERS: Architecture

53

Classification engine

Novelty detector

Re-labeler

Re-trainer

54 of 81

Classification Engine

54

Recognizes emotion and generates explanation

Any multimodal classification model such as EMERSK can be used

55 of 81

Novelty Detector

55

GateKeeper

Regenerator

Discriminator

56 of 81

GateKeeper

56

GateKeeper

Receives modular outputs

Marks novelty

57 of 81

Regenerator and Discriminator

57

Autoencoder based regenerator and CNN based discriminator

Encoder compresses the sample and Decoder reconstructs the sample

Reconstruction training with reconstruction loss

Adversarial training of R+D with minimax loss function

 

 

58 of 81

Novelty Handling

58

    • Periodic
    • Label correction
    • New class creation

Relabeling

    • Periodic
    • Human in the loop

Retraining

59 of 81

Evaluation Setup: Similar works

59

Name

Method

Limitations

Continuous learning loop?

Pix CNN [64]

Gated PixelCNN

Low accuracy

No

AnoGAN [65]

GAN+ coupled mapping

Inefficient

No

DSVDD [66]

Kernel-based one-class classification + minimum volume estimation

Limited generalization

No

60 of 81

Experimental Results and Findings: Novelty Detection

60

60

Research question: Is our detector reliable?

60

Findings: Our novelty detector offers superior detection capability.

Results on MNIST dataset

  • Trained in MNIST dataset
  • 50% samples of the test set are novelty
  • Example: Digit “1” to “8” inliers and “9” novelty

MNIST samples

61 of 81

Experimental Results and Findings: Retraining

61

61

Research question: Does our system improves performance?

61

Findings: In situations with frequent novel samples, our method can adapt and offer improved performance!

  • First cycle (no novelty): Regular model with 20% novelty samples in the test set---🡪
  • Samples detected using the detector and model retrained
  • Second cycle (with novelty): Updated model--🡪

Degraded accuracy!!

Improvement!!

Results on FER-2013 dataset with 20%

Novelty samples

62 of 81

Research Contribution of CoNERS

62

Formalization of the novelty in

the automatic emotion recognition task

An adversarially trained auto-encoder based detector for novelty detection

A system that addresses novelty in a continuous learning manner for emotion recognition

63 of 81

Conclusions and Future Works

63

64 of 81

Conclusion

64

Proposed SAFER, a novel system for emotion recognition from facial expressions

Proposed EMERSK, a multimodal emotion recognition for additional reliability and explainable output

Proposed CoNERS, a novelty-aware emotion recognition system for real world situation

65 of 81

Future Works

65

Enhanced multimodal fusion

Anxiety and depression detection

Fast and light-weight system building for real time operation

Addressing the ethical, security and privacy issues

66 of 81

Acknowledgement

66

Special Thanks

67 of 81

Thanks Everyone!

67

68 of 81

Extra Slides

68

69 of 81

Publications

69

Mijanur Rahaman Palash, Bharat Bhargava, SAFER: Situation Aware Facial Emotion Recognition, under review, Elsevier Artificial Intelligence 2023

Mijanur Rahaman Palash, Bharat Bhargava, EMERSK -Explainable Multimodal Emotion Recognition with Situational Knowledge, accepted for publication with minor revisions in IEEE Transaction on Multimedia 2023

Mijanur Rahaman Palash, Bharat Bhargava, Continuous Learning Based Novelty Aware Emotion Recognition System, AAAI Spring Symposium 2022

Mijanur Rahaman Palash, Voicu Popescu, Amit Sheoran and Sonia Fahmy, CoRE- Non-Linear 3D Sampling for Robust 360 Degree Video Streaming IEEE INFOCOM 2021

Mijanur Rahaman Palash, Bharat Bhargava, CoNERS: Continuous Learning Based Novelty Aware Emotion Recognition, under review, IEEE Transaction on Neural Networks and Learning Systems 2023

70 of 81

Reference:

70

[1] "The Seven Universal Emotions We Wear on Our Face," CBC, accessed on [date], available at: [https://www.cbc.ca/natureofthings/features/theseven-universal-emotions-we-wear-on-our-face].

[2] K. Patel, D. Mehta, C. Mistry, et al., "Facial Sentiment Analysis Using AI Techniques: State-of-the-Art, Taxonomies, and Challenges," IEEE Access, vol. 8, pp. 90,495-90,519, 2020.

[3] "Reading Facial Expressions of Emotion," American Psychological Association, accessed on [date], available at: [https://www.apa.org/science/about/psa/2011/05/facialexpressions.].

[4] T. Mittal, P. Guhan, U. Bhattacharya, R. Chandra, A. Bera, and D. Manocha, "Emoticon: Context-Aware Multimodal Emotion Recognition Using Frege's Principle," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 14,234-14,243.

[5] R. Kosti, J. M. Alvarez, A. Recasens, and A. Lapedriza, "Context-Based Emotion Recognition Using Emotic Dataset," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 11, pp. 2,755-2,766, 2019.

[6] J. Lee, S. Kim, S. Kim, J. Park, and K. Sohn, "Context-Aware Emotion Recognition Networks," in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 10,143-10,152.

[7] S. Knobloch-Westerwick, J. Abdallah, and A. C. Billings, "The Football Boost? Testing Three Models on Impacts on Sports Spectators' Self-Esteem," Communication & Sport, vol. 8, no. 2, pp. 236-261, 2020.

[8] J. Jayalekshmi and T. Mathew, "Facial Expression Recognition and Emotion Classification System for Sentiment Analysis," in 2017 International Conference on Networks & Advances in Computational Technologies (NetACT), IEEE, 2017, pp. 1-8.

[9] N. B. Kar, K. S. Babu, A. K. Sangaiah, and S. Bakshi, "Face Expression Recognition System Based on Ripplet Transform Type II and Least Square SVM," Multimedia Tools and Applications, vol. 78, no. 4, pp. 4,789-4,812, 2019.

[10] H. M. Shah, A. Dinesh, and T. S. Sharmila, "Analysis of Facial Landmark Features to Determine the Best Subset for Finding Face Orientation," in 2019 International Conference on Computational Intelligence in Data Science (ICCIDS), IEEE, 2019, pp. 1-4.

[11] V. Bazarevsky, Y. Kartynnik, A. Vakunov, K. Raveendran, and M. Grundmann, "BlazeFace: Sub-Millisecond Neural Face Detection on Mobile GPUs," arXiv preprint arXiv:1907.05047, 2019.

[12] R. S. Jadhav and P. Ghadekar, "Content-Based Facial Emotion Recognition Model Using Machine Learning Algorithm," in 2018 International Conference on Advanced Computation and Telecommunication (ICACAT), IEEE, 2018, pp. 1-5.

[13] "FER-2013 Learn Facial Expressions from an Image," Kaggle, accessed on [date], available at: [https://www.kaggle.com/msambare/fer2013].

[14] S. Datta, D. Sen, and R. Balasubramanian, "Integrating Geometric and Textural Features for Facial Emotion Classification Using SVM Frameworks," in Proceedings of the International Conference on Computer Vision and Image Processing, Springer, 2017, pp. 619-628.

[15] A. R. Kurup, M. Ajith, and M. M. Ramón, "Semi-Supervised Facial Expression Recognition Using Reduced Spatial Features and Deep Belief Networks," Neurocomputing, vol. 367, pp. 188-197, 2019.

[16] Y. Gan, J. Chen, and L. Xu, "Facial Expression Recognition Boosted by Soft Label with a Diverse Ensemble," Pattern Recognition Letters, vol. 125, pp. 105-112, 2019. "[date]" and "[URL]" should be replaced with the specific date and URL of access.

71 of 81

Reference:

71

[17] P. Dhankhar, "ResNet-50 and VGG-16 for Recognizing Facial Emotions," International Journal of Innovations in Engineering and Technology (IJIET), vol. 13, no. 4, pp. 126-130, 2019.

[18] A. P. Fard and M. H. Mahoor, "AD-Corre: Adaptive Correlation-Based Loss for Facial Expression Recognition in the Wild," IEEE Access, vol. 10, pp. 26,756-26,768, 2022.

[19] A. H. Farzaneh and X. Qi, "Facial Expression Recognition in the Wild via Deep Attentive Center Loss," in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2021, pp. 2402-2411.

[20] Y. Li, J. Zeng, S. Shan, and X. Chen, "Occlusion-Aware Facial Expression Recognition Using CNN with Attention Mechanism," IEEE Transactions on Image Processing, vol. 28, no. 5, pp. 2439-2450, 2018.

[21] K. Wang, X. Peng, J. Yang, D. Meng, and Y. Qiao, "Region Attention Networks for Pose and Occlusion Robust Facial Expression Recognition," IEEE Transactions on Image Processing, vol. 29, pp. 4057-4069, 2020.

[22] J. She, Y. Hu, H. Shi, J. Wang, Q. Shen, and T. Mei, "Dive into Ambiguity: Latent Distribution Mining and Pairwise Uncertainty Estimation for Facial Expression Recognition," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 6248-6257.

[23] B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba, "Places: A 10 Million Image Database for Scene Recognition," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, no. 6, pp. 1452-1464, 2017.

[24] D. Borth, R. Ji, T. Chen, T. Breuel, and S.-F. Chang, "Large-Scale Visual Sentiment Ontology and Detectors Using Adjective Noun Pairs," in Proceedings of the 21st ACM International Conference on Multimedia, 2013, pp. 223-232.

[25] A. Mollahosseini, B. Hasani, and M. H. Mahoor, "AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild," IEEE Transactions on Affective Computing, vol. 10, no. 1, pp. 18-31, 2017.

[26] S. Li, W. Deng, and J. Du, "Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2852-2861.

[27] P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, "The Extended Cohn-Kanade Dataset (CK+): A Complete Dataset for Action Unit and Emotion-Specified Expression," in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition - Workshops, 2010, pp. 94-101. doi: [DOI].

[28] J. Lee, S. Kim, S. Kim, J. Park, and K. Sohn, "Context-Aware Emotion Recognition Networks," in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 10,143-10,152.

[29] H. Gunes and M. Piccardi, "Bi-Modal Emotion Recognition from Expressive Face and Body Gestures," Journal of Network and Computer Applications, vol. 30, no. 4, pp. 1334-1345, 2007.

[30] P. Ekman, "An Argument for Basic Emotions," Cognition & Emotion, vol. 6, no. 3-4, pp. 169-200, 1992.

[31] K. He, X. Zhang, S. Ren, and J. Sun, "Deep Residual Learning for Image Recognition," 2015. arXiv: [arXiv link].

[32] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, "ImageNet: A Large-Scale Hierarchical Image Database," in 2009 IEEE Conference on Computer Vision and Pattern Recognition, IEEE, 2009, pp. 248-255.

[33] W. Li, X. Dong, and Y. Wang, "Human Emotion Recognition with Relational Region-Level Analysis," IEEE Transactions on Affective Computing, 2021.

[34] G. Wen, Z. Hou, H. Li, D. Li, L. Jiang, and E. Xun, "Ensemble of Deep Neural Networks with Probability-Based Fusion for Facial Expression Recognition," Cognitive Computation, vol. 9, no. 5, pp. 597-610, 2017.

72 of 81

Reference:

72

[35] A. Renda, M. Barsacchi, A. Bechini, and F. Marcelloni, "Comparing Ensemble Strategies for Deep Learning: An Application to Facial Expression Recognition," Expert Systems with Applications, vol. 136, pp. 1-11, 2019.

[36] D. Zeng, Z. Lin, X. Yan, Y. Liu, F. Wang, and B. Tang, "Face2Exp: Combating Data Biases for Facial Expression Recognition," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 20,291-20,300.

[37] K. Wang, X. Peng, J. Yang, S. Lu, and Y. Qiao, "Suppressing Uncertainties for Large-Scale Facial Expression Recognition," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 6897-6906.

[38] CDC, accessed on [date], available at: [URL].

[39] J. Chakraborty, S. Majumder, and T. Menzies, "Bias in Machine Learning Software: Why? How? What to Do?" in Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2021, pp. 429-440.

[40] Defi Dataset, accessed on [date], available at: [URL].

[41] A. Mollahosseini, D. Chan, and M. H. Mahoor, "Going Deeper in Facial Expression Recognition Using Deep Neural Networks," in 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), IEEE, 2016, pp. 1-10.

[42] T. Randhavane, U. Bhattacharya, P. Kabra, et al., "Learning Gait Emotions Using Affective and Deep Features," in Proceedings of the 15th ACM SIGGRAPH Conference on Motion, Interaction and Games, 2022, pp. 1-10.

[43] S. K. D'mello and A. Graesser, "Multimodal Semi-Automated Affect Detection from Conversational Cues, Gross Body Language, and Facial Features," User Modeling and User-Adapted Interaction, vol. 20, no. 2, pp. 147-187, 2010.

[44] U. Bhattacharya, T. Mittal, R. Chandra, T. Randhavane, A. Bera, and D. Manocha, "STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, 2020, pp. 1342-1350.

[45] H. Liu, H. Cai, Q. Lin, X. Li, and H. Xiao, "Adaptive Multilayer Perceptual Attention Network for Facial Expression Recognition," IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 9, pp. 6253-6266.

[46] J. L. Joseph and S. P. Mathew, "Facial Expression Recognition for the Blind Using Deep Learning," in 2021 IEEE 4th International Conference on Computing, Power and Communication Technologies (GUCON), 2021, pp. 1-5. doi: [DOI].

[47] P. Tahghighi, A. Koochari, and M. Jalali, "Deformable Convolutional LSTM for Human Body Emotion Recognition," in Pattern Recognition. ICPR International Workshops and Challenges: Virtual Event, January 10-15, 2021, Proceedings, Part III, Springer, 2021, pp. 741-747.

[48] P. D. Marrero Fernandez, F. A. Guerrero Pena, T. Ren, and A. Cunha, "FerAtt: Facial Expression Recognition with Attention Net," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2019, pp. 0-0.

[49] K. Sikka, K. Dykstra, S. Sathyanarayana, G. Littlewort, and M. Bartlett, "Multiple Kernel Learning for Emotion Recognition in the Wild," in Proceedings of the 15th ACM on International Conference on Multimodal Interaction, 2013, pp. 517-524.

[50] K. R. Scherer and H. Ellgring, "Multimodal Expression of Emotion: Affect Programs or Componential Appraisal Patterns?" Emotion, vol. 7, no. 1, p. 158, 2007.

[51] G. Castellano, M. Mortillaro, A. Camurri, G. Volpe, and K. Scherer, "Automated Analysis of Body Movement in Emotionally Expressive Piano Performances," Music Perception, vol. 26, no. 2, pp. 103-119, 2008.

[52] L. Chen, M. Li, M. Wu, W. Pedrycz, and K. Hirota, "Coupled Multimodal Emotional Feature Analysis Based on Broad-Deep Fusion Networks in Human-Robot Interaction," IEEE Transactions on Neural Networks and Learning Systems, 2023.

73 of 81

Reference:

73

[53] S. Poria, I. Chaturvedi, E. Cambria, and A. Hussain, "Convolutional MKL Based Multimodal Emotion Recognition and Sentiment Analysis," in 2016 IEEE 16th International Conference on Data Mining (ICDM), IEEE, 2016, pp. 439-448.

[54] B. Sun, S. Cao, J. He, and L. Yu, "Affect Recognition from Facial Movements and Body Gestures by Hierarchical Deep Spatio-Temporal Features and Fusion Strategy," Neural Networks, vol. 105, pp. 36-51, 2018.

[55] M. Li, L. Chen, M. Wu, W. Pedrycz, and K. Hirota, "Multimodal Information-Based Broad and Deep Learning Model for Emotion Understanding," in 2021 40th Chinese Control Conference (CCC), IEEE, 2021, pp. 7410-7414.

[56] T. Mittal, A. Bera, and D. Manocha, "Multimodal and Context-Aware Emotion Perception Model with Multiplicative Fusion," IEEE MultiMedia, vol. 28, no. 2, pp. 67-75, 2021.

[57] Y. Bhatia, A. H. Bari, and M. Gavrilova, "A LSTM-Based Approach for Gait Emotion Recognition," in 2021 IEEE 20th International Conference on Cognitive Informatics & Cognitive Computing (ICCI*CC), IEEE, 2021, pp. 214-221.

[58] T. Mittal, U. Bhattacharya, R. Chandra, A. Bera, and D. Manocha, "M3ER: Multiplicative Multimodal Emotion Recognition Using Facial, Textual, and Speech Cues," in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, 2020, pp. 1359-1367.

[59] K. Wang, X. Zeng, J. Yang, et al., "Cascade Attention Networks for Group Emotion Recognition with Face, Body and Image Cues," in Proceedings of the 20th ACM International Conference on Multimodal Interaction, 2018, pp. 640-645.

[60] T. Gedeon, A. Dhall, J. Joshi, J. Hoey, R. Goecke, and S. Ghosh, "From Individual to Group-Level Emotion Recognition: EmotiW 5.0," 2021.

[61] E. A. Veltmeijer, C. Gerritsen, and K. Hindriks, "Automatic Emotion Recognition for Groups: A Review," IEEE Transactions on Affective Computing, 2021.

[62] V. Bazarevsky, I. Grishchenko, K. Raveendran, T. Zhu, F. Zhang, and M. Grundmann, "BlazePose: On-Device Real-Time Body Pose Tracking," arXiv preprint arXiv:2006.10204, 2020. [64] V. Badrinarayanan, A. Kendall, and R. Cipolla, "SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 12, pp. 2481-2495, 2017.

[63] R. Santhoshkumar and M. Kalaiselvi Geetha, "Vision-Based Human Emotion Recognition Using HOG-KLT Feature," in Proceedings of First International Conference on Computing, Communications, and Cyber-Security (IC4S 2019), Springer, 2020, pp. 261-272.

[64] A. Van den Oord, N. Kalchbrenner, L. Espeholt, O. Vinyals, A. Graves, et al., “Conditional image generation with pixelcnn decoders,” Advances in neural information processing systems, vol. 29, 2016.

[65] T. Schlegl, P. Seeböck, S. M. Waldstein, U. Schmidt-Erfurth, and G. Langs, “Unsupervised anomaly detection with generative adversarial networks to guide marker discovery,” in Information Processing in Medical Imaging: 25th International Conference, IPMI 2017,

Boone, NC, USA, June 25-30, 2017, Proceedings, Springer, 2017, pp. 146–157.

[66] L. Ruff, R. Vandermeulen, N. Goernitz, et al., “Deep one-class classification,” in International conference on machine learning, PMLR, 2018, pp. 4393–4402.

74 of 81

Data Bias

  • How do we deal with bias
    • Machine learning algorithms can discriminate based on classes like race and gender
    • A good model is dependent on a good dataset and without proper care a dataset can lack diversity
    • Biased dataset will perform poorly with minority:
      • If most of the samples are white males, the model will fail for women and people of color
    • Researchers (Buolamwini et. al.) showed:
      • Three commercially released facial-analysis programs from major technology companies demonstrate both skin-type and gender biases
      • Error rates in determining the gender of light-skinned men were never worse than 0.8 percent
      • For darker-skinned women, more than 20 percent in one case and more than 34 percent in the other two

74

75 of 81

Data Bias

  • How bias is introduced in ER
    • Keyword searching in google is a popular method of collecting visual (image and video) data
    • In our search with keyword “angry face”- 85% of the acceptable images appeared are male
    • This pattern holds for other generic keywords like ”sad people”, ”happy human” etc.
    • Therefore, a dataset prepared by collecting results from these types of keyword search results in bias
    • Same applies to the volunteer choice for creating an acted dataset
    • Without careful selection of people from multiple genders and ethnic backgrounds, dataset bias can be easily incorporated into the model

75

76 of 81

Data Bias

  • Bias reduction plan
    • Better representation of minority groups by using specific keywords:
      • Using both “happy man face” and “happy woman face” instead of “happy face” keyword
    • Choosing volunteers from diverse background
    • To produce new ML models which provide higher importance on less represented data samples
    • Data augmentation
    • Data cleaning algorithms
    • Transfer learning

76

77 of 81

Data Bias

  • Bias in the ER datasets
    • Widely used ER dataset FER-2013 is an example of keyword search bias

77

78 of 81

CNN

  • Input Image: 226x226x3

  • Convolutional Layer 1:
  • Filter Size: 2x2
  • Stride: 1
  • Padding: 0
  • Output Dimension: 225x225x3
  • Activation: ReLU
  • Max Pooling Layer 1:
  • Pooling Size: 2x2
  • Output Dimension: 112x112x3

  • Convolutional Layer 2:
  • Filter Size: 2x2
  • Stride: 1
  • Padding: 0
  • Output Dimension: 111x111x3
  • Activation: ReLU
  • Max Pooling Layer 2:
  • Pooling Size: 2x2
  • Output Dimension: 55x55x3

78

79 of 81

CNN

  • Convolutional Layer 3:
  • Filter Size: 2x2
  • Stride: 1
  • Padding: 0
  • Output Dimension: 54x54x3
  • Activation: ReLU
  • Max Pooling Layer 3:
  • Pooling Size: 2x2
  • Output Dimension: 27x27x3

  • Fully Connected Layer:
  • Input Dimension: 27x27x3
  • Output Dimension: 256

79

80 of 81

Emotion Recognition Use Cases

80

    • Identify mental health issue to prevent school shooting

Public Safety

    • Identify suspicious behavior and criminal intent

Law Enforcement

    • Detect medical conditions such as depression

Healthcare

    • Trigger alarm for extreme emotional state (anger, fear etc.) of the driver

Autonomous Car

    • Adjust the gameplay based on the comfort level of the player

Interactive Gaming

81 of 81

81