1 of 31

BEAUTY MOMENT SYNTHESIS

NGA VU, PHUC NGUYEN, THINH TRAN – CAPSTONE PROJECT LHP.DL4AI.02.2022

SPECIAL THANKS TO OUR ADVISOR/MENTOR: MR. TIEP NGUYEN – PHD, LECTURER AT UIT VNU-HCM AND VIETAI

2 of 31

IN THIS PRESENTATIONS

  • PROJECT INTRODUCTION
  • BASELINE INTRODUCTION
  • DETAILED BASELINE
  • DATASET
  • MODULE 1: FACE DETECTION AND RECOGNITION
  • MODULE 2: FACE IMAGE QUALITY ASSESSMENT
  • MODULE 3: SMILE SCORE EVALUATION
  • MODULE 4: VIDEO SYNTHESIS
  • FINAL PRODUCT
  • FUTURE DEVELOPMENTS
  • DEPLOYMENT AND DEMO

3 of 31

PROJECT INTRODUCTION

  • Our projects aim to apply multiple Deep Learning models to create a final pipeline to find some of the most “beautiful” images in our given image set and synthesize them into a video.
  • Inputs: 2 image folders
    • Anchor images of people that we need to recognize in our image set.
    • An image set, in which we will find the “most beautiful” images of the target people in our anchor folder.
  • Output: synthesized video (.mp4)

4 of 31

PROJECT TIMELINE

March 2022

Project brainstorm and reproduce different modules

April 2022

Combines modules

May 2022

Optimizations and evaluations

June 2022

Products demo and presentation

Future work

More optimizations

More features

5 of 31

OVERALL BASELINE

Face recognition from targets in anchor folder

Face image quality assessment (FIQA) from the extracted bounding boxes from face recognition modules

Synthesize the qualified images into video with transitions and animations

Smile score: evaluate and determine if the targets are smiling

01

03

02

04

Anchor

Image Set

6 of 31

DETAILED BASELINE

FACE DETECTION

ANCHOR IMAGES OF 6 PEOPLE

INPUT IMAGES

FACE RECOGNITION

CROPPED FACES

FACE ALIGNMENT

KNN

PRETRAINED MODEL

Bounding boxes with id

Filter disqualified bounding boxes

Ranking images based on smile scores

Quality scores

Smile scores

Video synthesis

7 of 31

DATASET

  • Through the process, we tried and tested our modules on various data sets:
    • Emotion Detection dataset on Kaggle (this link): training and testing for smile score module
    • Wider-face data set: you can read our EDA in this link
  • This data set was used for testing our bounding box cutting algorithms, face detection module, and face image quality assessment module
    • Later, we implemented these techniques to test our baseline, in which we used a private data set, provided by Mr. Tiep Nguyen

8 of 31

OUR PRIVATE DATASET (PROVIDED BY MR. TIEP NGUYEN)

  • Includes 623 images provided by Mr. Tiep Nguyen.
  • The images are taken at different times of the day: morning, afternoon, and evening.
  • Our anchor set includes pictures of six people at the ceremony.

Sample images from the Image Set

Sample images from the Anchor Set

9 of 31

OUR DATASET STRUCTURE

→ Inside the input folder, you can have as many as sub-folders - it does not matter, as we are using recursion in order to read all images inside your input directory. The same rule applies in the anchor set.

10 of 31

MODULE 1: FACE DETECTION AND RECOGNITION

  • Task: detect and recognize faces of the target people in all images
  • Experience: 2 datasets - anchor dataset and image set
  • Performance: Average IoU score - 0.56 ; mAP - 0.5 ; F1 average score - 0.65; ROC AUC average score - 0.82
  • Function space: MTCNN (detect human faces), a pretrained InceptionResnetV1 (provided by Facenet) (extract detected human faces into 512-dimensional vectors), and K-nearest neighbor (recognize people using a cosine-similarity threshold )
  • Algorithm: MTCNN, InceptionResnetV1 (optimized with Adam, pre-trained on VGGFace 2 dataset with 99.65 % accuracy), K-nearest neighbors

11 of 31

Face Alignment Examples

12 of 31

Face Recognition Examples

Vũ Hải Quân

Võ Văn Khang

Vũ Hải Quân

13 of 31

When Face Recognition module goes wrong

14 of 31

MODULE 2: FACE IMAGE QUALITY ASSESSMENT

  • You can read more about SDD-FIQA (Unsupervised Face Image Quality Assessment with Similarity Distribution Distance) at this GitHub repo and read its paper here.
  • Applied pre-trained model (about 30M params) in PyTorch to predict face images’ quality score and remove unqualified images with no bounding boxes qualified for the threshold.

15 of 31

MODULE 2: FACE IMAGE QUALITY ASSESSMENT

  • Task: create quality score of face bounding boxes detected by the previous module
  • Experience: MS-Celeb-1M data set
  • Performance (on our data set):
  • Function space: scalar outputs of ResNet50
  • Algorithm: Adam

Link to paper: https://arxiv.org/pdf/2103.05977.pdf - SDD-FIQA (CVPR-2021)

Precision

Recall

F1-score

Support

Good

0.38

0.59

0.46

96

Bad

0.78

0.59

0.67

232

Accuracy

0.59

328

Macro avg

0.58

0.59

0.57

328

Weighted avg

0.66

0.59

0.61

328

SDD-FIQA classification report on our data set

16 of 31

HOW SDD-FIQA CREATE FACE IMAGES’ QUALITY SCORE?

  • Let’s see some examples of scores generated by SDD-FIQA model:

41.908226

35.694916

44.77323

17 of 31

HOW DIFFERENT SITUATION AFFECTS FIQA SCORES?

  • Let’s see some examples of scores generated by SDD-FIQA model:

45.356842

0.0349

1.2805

42.950836

18 of 31

HOW DIFFERENT SITUATION AFFECTS FIQA SCORES?

41.908226

28.20411

41.72763

32.48454

31.134018

27.354168

32.721783

32.721783

19 of 31

HOW DIFFERENT SITUATION AFFECTS FIQA SCORES?

35.25

33.97

43.53

34.15

44.27

28.98

46.19

30.00

20 of 31

HOW DIFFERENT SITUATION AFFECTS FIQA SCORES?

39.73

33.02

43.53

43.32

43.00

21.51

45.67

21 of 31

MODULE 3: SMILE SCORE ASSESSMENT

  • Task: create smile score of face bounding boxes detected by the previous module
  • Experience: Fer2013 dataset
  • Performance: (on our data set):

  • Function space: scalar outputs of CNN model
  • Algorithm: Adam

22 of 31

23 of 31

Overview of the dataset Fer2013

Angry

Disgust

Fear

Happy

Neutral

sad

Surprise

24 of 31

HOW DEEPFACE CREATE FACE IMAGES’ SMILE SCORE?

  • Let’s see some examples of scores generated by Deepface emotion module:

0.607

97.07

17.07

25 of 31

Deepface architecture

26 of 31

MODULE 4: VIDEO SYNTHESIS

TRANSITIONS

ANIMATIONS

Cover animation

Comb animation

Push animation

Uncover animation

Split animation

Fade animation

Zoom animation

Rotate animation

are coded from scratch and randomly selected into the final video.

27 of 31

MODULE 4: VIDEO SYNTHESIS

  • Example of testing and optimizing cover animation:

28 of 31

PROBLEMS AND SOLUTIONS

PROBLEMS

  • Too many frames to render – out of memory

SOLUTIONS

  • Create tmp.mp4 for later video merging

MORE ON OUR SOLUTION:

OUR BASELINE

SMALLER VIDEO FILES

CONCATENATE

FINAL VIDEO

DELETE TEMPORARY FILES

29 of 31

PROBLEMS AND SOLUTIONS

PROBLEMS

  • The shape of bounding box to zoom in does not match the image’s aspect ratios

SOLUTIONS

  • Create algorithm to create new bounding boxes that match the images’ aspect ratios

our algorithm

30 of 31

FINAL PRODUCT

These are the best moments of Mr. Vu Hai Quan throughout more than 600 images from our image set.

Bounding boxes are annotated for debugging purposes.

31 of 31

FUTURE DEVELOPMENTS

WHAT WE HAVE DONE

  • Finished building a comprehensive baseline for our project.
  • Finished building an evaluation pipeline + ground-truth dataset to evaluate our modules.
  • Finished building a basic demo web for our project.

OUR PLANS FOR THE FUTURE

  • Fine-tune and re-train our models on the benchmark dataset.
  • Optimize our modules speed.
  • Add more effects to our video generator.
  • Insert music based on the scenes in the pictures for our video.
  • Add option to extract the most beautiful frames from a video and synthesize it into an output video.