1 of 17

XI INTERNATIONAL CONFERENCE

“INFORMATION TECHNOLOGY AND IMPLEMENTATION” (IT&I-2024)

Optimal size reduction methodology for YOLO-based underwater object detectors based on knowledge distillation

​

Victor SINEGLAZOV 1, 2, Mykhailo SAVCHENKO 1, Michael Z. ZGUROVSKY 1

​

1 National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”, 37, Prospect Beresteiskyi, Kyiv, 03056, Ukraine�2 National Aviation University, 1, Prospect Liubomyra Huzara, Kyiv, 03058, Ukraine

​

2 of 17

Autonomous underwater vehicles (AUVs)

AUVs are unmanned devices, widely used in water resources research. They allow:

​

  • working completely autonomously and do not require human assistance;
  • operating at considerable depths and distances;
  • obtaining the necessary visual data and information from sensors cost-effectively and quickly.

Figure 1. Autonomous underwater vehicle examples.

3 of 17

Applications of AUVs

Enabling AUVs with object detection technologies allows to use it in variety of applications, including, but not limited to:

​

  • Biology: exploring biodiversity of seas and oceans.
  • Archeology: looking for underwater historical objects.
  • Oil and gas: exploring ocean floor to find optimal routes.
  • Ecology: pollution monitoring.
  • Military: demining and security monitoring.
  • Rescue services: exploring ship wreckages.

Figure 2. US Navy AUV before launch.

4 of 17

Problems of creating intellectual object detection systems for AUVs

Two main problems of creating a deep-learning based intellectual system for object detection of underwater objects, are the following:

​

  • A proposed object detection neural network should provide sufficient accuracy on images taken in underwater environment with visibility issues, small and densely packed objects.
  • Apart from being accurate, this network should be small and efficient to be deployed on edge hardware, carried by AUV, which has limited performance and battery life.

Figure 3. Sample images, taken by autonomous underwater vehicle.

5 of 17

The purpose of this study

Main purpose of our study is to create an optimal size-reduction methodology for underwater object detectors, based on YOLO network architecture. The resulting methodology should allow:

​

  • developing an object detection neural network, which will work correctly with images taken in underwater environments;
  • deploying the resulting neural network on low-powered and low-compute hardware, which can be installed as an external module for AUV.

​

Figure 4. Sample of the desired output from proposed network.

6 of 17

Problem statement

 

7 of 17

Dataset

  • Underwater Target Detection Algorithm Challenge 2020 dataset (UTDAC2020) has been used.
  • The dataset features 5168 images in 4 classes: echinus, holothurian, scallop and starfish.
  • The classes in the dataset are imbalanced, with 23 648 objects in echinus class, and only 7 838 and 5 245 objects in holothurian and scallop classes respectively.
  • 1292 images are used to validate model performance.

Figure 5. Example of an image from the dataset.

8 of 17

Proposed approach

The methodology, proposed in this study, includes the following methods:

​

  • Developing knowledge distillation algorithm, which enables transferring knowledge from teacher to student model, both based on YOLO architecture.
  • Building light-weight student model topology, which utilizes computationally effective convolutional layers to reach the desired size and parameter count reduction effect.

Figure 6. Overall structure of neural network, built upon our methodology.

9 of 17

Developing knowledge distillation algorithm for YOLO

To enable knowledge transfer from larger teacher model into light-weight student model, regression and classification components of total loss function have been modified to include distillation loss with additional weighting coefficient added to avoid learning collapse due to student model fully mimicking teacher model outputs. Additionally, temperature coefficient 𝜏 with decay strategy has been used to control the softening of logits for the classification component, allowing to regulate the amount of knowledge being distilled from a larger model by a student.

Figure 6. The proposed algorithm.

10 of 17

Developing knowledge distillation algorithm for YOLO

 

11 of 17

Building light-weight student model topology

To ensure optimal performance of a resulting distilled model, the student model should meet the size and computational efficiency requirement. An approach used in this paper involved light-weighting the feature extraction (backbone) layers of YOLO object detector, to reduce the number of expensive convolutional operations, which contribute a lot to a total parameter count. Feature aggregation (neck) and final output layers (head) from original YOLOv8 architecture were reused, as adding additional blocks to these parts of the network would increase the parameter count, and extra light-weighting would introduce more differences between student and teacher model, which could harm the distillation performance.

​

Following backbone architectures were proposed:

  • based on convolutional operations fouud in GhostNet architecture;
  • based on partial convolutions, derived from FasterNet architecture.

12 of 17

Building light-weight student model topology — GhostNet based

 

Figure 7. GhostNet approach to convolution, schematically.

13 of 17

Building light-weight student model topology — FasterNet based

FasterNet applies different approach to reduce the computational complexity and decrease latency of convolutional operations, based on PConv procedure, which applies convolutional operation only on a part of input channels for spatial feature extraction and leaves remaining channels as is. Then, PConv is followed by series of pointwise convolution to reuse the information from all channels in an efficient way.

​

Figure 8. FasterNet approach to convolution called PConv, schematically.

14 of 17

Metrics and experimental setup

Total of five metrics have been used to test the model, with mAp and mAp50 representing the object detection accuracy of neural network. Size, parameter count and FLOPs are also measured as performance metrics to evaluate the computational efficiency of proposed approach.

​

  • mAp50 — mean average precision (mAp) at intersection over union (IoU) of 0.50, defined as average precision for each class over number of classes.
  • mAp — mAp at IoU of 0.50:0.05:0.95.
  • Params — total number of model parameters.
  • FLOPs — performance metrics denoting number of floating-point operations per second.
  • Size — model size in megabytes.

​

The machine used for experiment is equipped with Intel Core i5-13600K processor, NVIDIA A4000 GPU with 16GB VRAM. Software-wise, the test setup is running Ubuntu 20.04.6 LTS with Python 3.10.13, CUDA 12.1, and PyTorch 2.2.1. Optimizer used — SGD w/ momentum 0.937, initial learning rate 0.01, weight decay coefficient 0.005, distillation temperature 5, augmentations handled by Albumentations library.

15 of 17

Experiment results on UTDAC2020 dataset

For distilled models, YOLOv8l with DarkNet-53 backbone is used as a teacher model. Student models use YOLOv8s architecture with custom backbones, based on GhostNet and FasterNet, with both convolutional blocks and bottlenecks modified. Models using knowledge distillation are marked with ‘-dist’ suffix.

Method

Backbone

mAp

mAp50

Params (M)

FLOPs (G)

Size (Mb)

Faster R-CNN

ResNet50

44.51

80.93

41.14

63.3

~

RetinaNet

ResNet50

43.93

80.42

36.17

52.6

~

FCOS

ResNet50

43.88

81.06

31.84

50.4

~

 

ResNet50

 

 

 

 

 

YOLOv8n

DarkNet-53

48.92

82.61

3

8.9

6

YOLOv8s

DarkNet-53

50.45

84.58

11.2

28.8

22

YOLOv8m

DarkNet-53

51.62

84.92

25.8

78.7

51

YOLOv8l

DarkNet-53

51.73

84.97

43.6

165.7

84

YOLOv8s

GhostNet

49.76

83.71

6

16.4

9

YOLOv8s

FasterNet

49.85

83.8

5.8

16

9

YOLOv8s-dist

GhostNet

50.62

84.7

6

16.4

9

YOLOv8s-dist

FasterNet

50.71

84.72

5.8

16

9

Table 1. Experimental results.

16 of 17

Examples of detections with proposed model

Figure 9. Examples of object detection obtained from YOLOv8s-dist model. Ground truth labels are on the left, proposed model detection results are on the right. Detection of targets at various scales and objects on complex backgrounds is handled correctly.

17 of 17

Conclusions

  • Our paper proposed a novel methodology to reduce the size of YOLO-based underwater object detectors. Knowledge distillation algorithm with temperature decay strategy has been developed for object detection neural network, allowing to effectively train light-weight student model by transferring knowledge from teacher model of larger capacity.
  • Additionally, we built two light-weight YOLO topologies, derived from GhostNet and FasterNet approaches to convolution operation, which are suitable to be used as a student model in knowledge distillation tasks.
  • The proposed light-weight models are 40% more efficient in terms of size, compared to existing YOLOv8s model. After using our knowledge distillation algorithm, the performance of the student model is superior to original YOLOv8s in terms of accuracy (84.72% and 84.58%, respectively), while the model size is comparable to YOLOv8n, the smallest model among YOLO-based detectors.
  • The applications of our methodology include training efficient object detection neural network for integrated autonomous underwater vehicle hardware.