XI INTERNATIONAL CONFERENCE
“INFORMATION TECHNOLOGY AND IMPLEMENTATION” (IT&I-2024)
Optimal size reduction methodology for YOLO-based underwater object detectors based on knowledge distillation
Victor SINEGLAZOV 1, 2, Mykhailo SAVCHENKO 1, Michael Z. ZGUROVSKY 1
1 National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”, 37, Prospect Beresteiskyi, Kyiv, 03056, Ukraine�2 National Aviation University, 1, Prospect Liubomyra Huzara, Kyiv, 03058, Ukraine
Autonomous underwater vehicles (AUVs)
AUVs are unmanned devices, widely used in water resources research. They allow:
Figure 1. Autonomous underwater vehicle examples.
Applications of AUVs
Enabling AUVs with object detection technologies allows to use it in variety of applications, including, but not limited to:
Figure 2. US Navy AUV before launch.
Problems of creating intellectual object detection systems for AUVs
Two main problems of creating a deep-learning based intellectual system for object detection of underwater objects, are the following:
Figure 3. Sample images, taken by autonomous underwater vehicle.
The purpose of this study
Main purpose of our study is to create an optimal size-reduction methodology for underwater object detectors, based on YOLO network architecture. The resulting methodology should allow:
Figure 4. Sample of the desired output from proposed network.
Problem statement
Dataset
Figure 5. Example of an image from the dataset.
Proposed approach
The methodology, proposed in this study, includes the following methods:
Figure 6. Overall structure of neural network, built upon our methodology.
Developing knowledge distillation algorithm for YOLO
To enable knowledge transfer from larger teacher model into light-weight student model, regression and classification components of total loss function have been modified to include distillation loss with additional weighting coefficient added to avoid learning collapse due to student model fully mimicking teacher model outputs. Additionally, temperature coefficient 𝜏 with decay strategy has been used to control the softening of logits for the classification component, allowing to regulate the amount of knowledge being distilled from a larger model by a student.
Figure 6. The proposed algorithm.
Developing knowledge distillation algorithm for YOLO
Building light-weight student model topology
To ensure optimal performance of a resulting distilled model, the student model should meet the size and computational efficiency requirement. An approach used in this paper involved light-weighting the feature extraction (backbone) layers of YOLO object detector, to reduce the number of expensive convolutional operations, which contribute a lot to a total parameter count. Feature aggregation (neck) and final output layers (head) from original YOLOv8 architecture were reused, as adding additional blocks to these parts of the network would increase the parameter count, and extra light-weighting would introduce more differences between student and teacher model, which could harm the distillation performance.
Following backbone architectures were proposed:
Building light-weight student model topology — GhostNet based
Figure 7. GhostNet approach to convolution, schematically.
Building light-weight student model topology — FasterNet based
FasterNet applies different approach to reduce the computational complexity and decrease latency of convolutional operations, based on PConv procedure, which applies convolutional operation only on a part of input channels for spatial feature extraction and leaves remaining channels as is. Then, PConv is followed by series of pointwise convolution to reuse the information from all channels in an efficient way.
Figure 8. FasterNet approach to convolution called PConv, schematically.
Metrics and experimental setup
Total of five metrics have been used to test the model, with mAp and mAp50 representing the object detection accuracy of neural network. Size, parameter count and FLOPs are also measured as performance metrics to evaluate the computational efficiency of proposed approach.
The machine used for experiment is equipped with Intel Core i5-13600K processor, NVIDIA A4000 GPU with 16GB VRAM. Software-wise, the test setup is running Ubuntu 20.04.6 LTS with Python 3.10.13, CUDA 12.1, and PyTorch 2.2.1. Optimizer used — SGD w/ momentum 0.937, initial learning rate 0.01, weight decay coefficient 0.005, distillation temperature 5, augmentations handled by Albumentations library.
Experiment results on UTDAC2020 dataset
For distilled models, YOLOv8l with DarkNet-53 backbone is used as a teacher model. Student models use YOLOv8s architecture with custom backbones, based on GhostNet and FasterNet, with both convolutional blocks and bottlenecks modified. Models using knowledge distillation are marked with ‘-dist’ suffix.
Method | Backbone | mAp | mAp50 | Params (M) | FLOPs (G) | Size (Mb) |
Faster R-CNN | ResNet50 | 44.51 | 80.93 | 41.14 | 63.3 | ~ |
RetinaNet | ResNet50 | 43.93 | 80.42 | 36.17 | 52.6 | ~ |
FCOS | ResNet50 | 43.88 | 81.06 | 31.84 | 50.4 | ~ |
| ResNet50 |
|
|
|
|
|
YOLOv8n | DarkNet-53 | 48.92 | 82.61 | 3 | 8.9 | 6 |
YOLOv8s | DarkNet-53 | 50.45 | 84.58 | 11.2 | 28.8 | 22 |
YOLOv8m | DarkNet-53 | 51.62 | 84.92 | 25.8 | 78.7 | 51 |
YOLOv8l | DarkNet-53 | 51.73 | 84.97 | 43.6 | 165.7 | 84 |
YOLOv8s | GhostNet | 49.76 | 83.71 | 6 | 16.4 | 9 |
YOLOv8s | FasterNet | 49.85 | 83.8 | 5.8 | 16 | 9 |
YOLOv8s-dist | GhostNet | 50.62 | 84.7 | 6 | 16.4 | 9 |
YOLOv8s-dist | FasterNet | 50.71 | 84.72 | 5.8 | 16 | 9 |
Table 1. Experimental results.
Examples of detections with proposed model
Figure 9. Examples of object detection obtained from YOLOv8s-dist model. Ground truth labels are on the left, proposed model detection results are on the right. Detection of targets at various scales and objects on complex backgrounds is handled correctly.
Conclusions