1 of 55

21/12/2023

Полухін Андрій

Efficiency in AI:

Практичні поради з оптимізації ШІ

DOU AI Meetup, Kyiv

2 of 55

Про мене

    • ML Engineer at Data Science UA and Samba.TV.
    • Writing about AI
      • Telegram (t.me/eiaioi), Medium, Website.
    • Ph. D. Student at NTUU KPI, “Unmanned Aerial Vehicle Navigation using Deep Learning on Edge Devices”, NATO Science for Peace and Security project G6032.
    • Mentor at Projector Mentorship Platform.
    • Teacher of Machine Learning at Hillel IT School.

3 of 55

Agenda

    • Навіщо оптимізувати ШІ
      • Оптимізація витрат та Data Privacy
    • Як можна оптимізувати ШІ
      • PyTorch, ONNX, TensorRT, TensorFlow, OpenVINO
      • Model Architecture, Distillation, Pruning, Mixed Precision, Quantization
    • Приклади (2 шт.)
      • YOLOv8, x3 speed up in a few lines of code
      • Consumer PC GPT

4 of 55

5 of 55

Переваги

6 of 55

Переваги

    • Персональний помічник
    • Генерує ідеї
    • Швидко відповідає
    • Мультимовний

7 of 55

    • Персональний помічник
    • Генерує ідеї
    • Швидко відповідає
    • Мультимовний

Переваги

Недоліки

8 of 55

    • Персональний помічник
    • Генерує ідеї
    • Швидко відповідає
    • Мультимовний
    • Збирає ваші дані
    • Працює лише online
    • Закритий код
    • Обмежене Usage Policy

Переваги

Недоліки

9 of 55

Альтернативи ChatGPT

    • Claude 2 (Anthropic)
    • Bard, Gemini (Google)
    • Coral (Cohere)
    • Grok (X)
    • Ernie (Baidu)

10 of 55

Альтернативи ChatGPT

Недоліки

    • Claude 2 (Anthropic)
    • Bard, Gemini (Google)
    • Coral (Cohere)
    • Grok (X)
    • Ernie (Baidu)
    • Закритий код
    • Обмежене Usage Policy
    • Збирає ваші дані
    • Працює лише online

11 of 55

Open Source LLM

12 of 55

Open Source LLM

Недоліки

Переваги

13 of 55

Open Source LLM

Недоліки

Переваги

    • Працює offline
    • Відкритий код
    • Немає обмежень на Usage Policy
    • Можна дотренувати під свої дані
    • Немає витрат на кількість використаних токенів
    • Не збирає статистику та ваші дані

14 of 55

Open Source LLM

Недоліки

Переваги

    • Працює offline
    • Відкритий код
    • Немає обмежень на Usage Policy
    • Можна дотренувати під свої дані
    • Немає витрат на кількість використаних токенів
    • Не збирає статистику та ваші дані
    • Потрібні великі обчислювальні потужності (>>RAM та GPU)
    • Потребують дотренування під задачу
    • Розгортання, підтримка та правильна робота (RAG, ToT) потребують окремих спеціалістів та витрат

15 of 55

Open Source LLM

16 of 55

Neural Network Inference

17 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Neural Network Inference

18 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

ML Framework

Neural Network Inference

19 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Data

ML Framework

Pre Processing

Post Processing

Inference

Pipeline

Neural Network Inference

20 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Data

ML Framework

Pre Processing

Post Processing

Inference

Pipeline

Neural Network Inference

21 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Data

ML Framework

Pre Processing

Post Processing

Inference

Pipeline

Neural Network Inference

22 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Data

ML Framework

Pre Processing

Post Processing

Inference

Pipeline

Neural Network Inference

23 of 55

Виберіть фреймворк

24 of 55

Виберіть фреймворк

    • general:
      • PyTorch (Meta AI)
      • TensorFlow (Google)

25 of 55

Виберіть фреймворк

    • general:
      • PyTorch (Meta AI)
      • TensorFlow (Google)
    • mobile:
      • TF Lite (Google)
      • TF Edge TPU (Google)
      • ncnn (Tencent)
      • Core ML (Apple)

26 of 55

    • optimized:
      • ONNX (Linux Foundation)
      • TensorRT (NVIDIA)
      • OpenVINO (Intel)
      • JAX (Google)
      • PaddlePaddle (PaddlePaddle)
      • DeepSparse (Neural Magic)
      • Mojo (Modular)
      • vLLM (vLLM)

Виберіть фреймворк

    • mobile:
      • TF Lite (Google)
      • TF Edge TPU (Google)
      • ncnn (Tencent)
      • Core ML (Apple)
    • general:
      • PyTorch (Meta AI)
      • TensorFlow (Google)

27 of 55

28 of 55

Any neural network

e.g. YOLO, Llama2, Stable Diffusion, Whisper

Data

ML Framework

Pre Processing

Post Processing

Inference

Pipeline

Neural Network Inference

29 of 55

Model Architecture

    • ConvNets Optimizations
      • ❌ Convolution ✅ Depthwise Separable Convolution (MobileNet)
      • ❌ Convolution ✅ Group Convolution (e.g. ShuffleNet)
      • ✅ Use Squeeze & Excitation (e.g. SqueezeNext)
      • ✅ Use Fused Inverted Residual Block (EfficientNet v2)
    • LLM & Transformer Optimizations
    • General Optimizations
      • ✅ Apply Neural Architecture Search (e.g. FBNet)

30 of 55

Model Architecture

31 of 55

Model Architecture

32 of 55

Model Architecture

    • ConvNets Optimizations
      • ❌ Convolution ✅ Depthwise Separable Convolution (MobileNet)
      • ❌ Convolution ✅ Group Convolution (e.g. ShuffleNet)
      • ✅ Use Squeeze & Excitation (e.g. SqueezeNext)
      • ✅ Use Fused Inverted Residual Block (EfficientNet v2)
    • LLM & Transformer Optimizations
    • General Optimizations
      • ✅ Apply Neural Architecture Search (e.g. FBNet)

33 of 55

Model Distillation

The teacher-student framework for knowledge distillation

34 of 55

Model Distillation

pip3 install torchdistill

35 of 55

Model Pruning

Pruning: Before and After

36 of 55

Model Pruning

37 of 55

Model Pruning

from neural_compressor.training import prepare_pruning, WeightPruningConfigconfig = WeightPruningConfig(configs)

prepare_pruning(model, config, optimizer)

for epoch in range(num_train_epochs):

model.train()

for step, batch in enumerate(train_dataloader):

outputs = model(**batch)

loss = outputs.loss

loss.backward()

optimizer.step()

lr_scheduler.step()

model.zero_grad()

pip3 install neural-compressor

38 of 55

Half Precision fp16

Figure represents comparison of FP16 (half precision floating points)

and FP32 (single precision floating points).

39 of 55

Half Precision fp16

# PyTorch

import torch

model = model.half() # Convert model to half precision

input_data = input_data.half() # Convert input data to half precision

# TensorFlow

import tensorflow as tf

from tensorflow.keras.mixed_precision import experimental as mixed_precision

policy = mixed_precision.Policy('mixed_float16')

mixed_precision.set_policy(policy)

# TensorRT

import tensorrt as trt

builder = trt.Builder(TRT_LOGGER)

builder.fp16_mode = True

40 of 55

Model Quantization

8-bit Integer (INT8)

41 of 55

Model Quantization

Quantization technique

42 of 55

Model Quantization

Deploy Just in 4 Lines with MQBench

model = models.__dict__["resnet18"](pretrained=True)

model = prepare_by_platform(model, BackendType.Tensorrt)

enable_calibration(model) # turn on calibration

for i, batch in enumerate(data): # do forward procedures

...

enable_quantization(model) # turn on actually quantization

input_shape={'data': [10, 3, 224, 224]}

convert_deploy(model, backend, input_shape) # model export

43 of 55

int4

MLPerf v0.5 Inference results for data center server form factors and offline scenario retrieved from www.mlperf.org on Nov. 6, 2019 (Closed Inf-0.5-25 and Inf-0.5-27 for INT8, Open Inf-0.5-460 and Inf-0.5-462 for INT4). Per-processor performance is calculated by dividing the primary metric of total performance by number of accelerators reported. MLPerf name and logo are trademarks.

44 of 55

Приклади

45 of 55

YOLOv8, x3 speed up

46 of 55

YOLOv8, x3 speed up

Specifications:

    • YOLO v8 m
    • g5.2xlarge
    • batch=64
    • imgsz=1280x1280

47 of 55

YOLOv8, x3 speed up

YOLO v8 m

56 FPS

Specifications:

    • YOLO v8 m
    • g5.2xlarge
    • batch=64
    • imgsz=1280x1280

48 of 55

YOLOv8, x3 speed up

YOLO v8 m

56 FPS

Specifications:

    • YOLO v8 m
    • g5.2xlarge
    • batch=64
    • imgsz=1280x1280

YOLO v8 m

91 FPS

Half Precision

49 of 55

YOLOv8, x3 speed up

YOLO v8 m

56 FPS

Specifications:

    • YOLO v8 m
    • g5.2xlarge
    • batch=64
    • imgsz=1280x1280

YOLO v8 m

91 FPS

Half Precision

YOLO v8 m

57 FPS

TensorRT

50 of 55

YOLOv8, x3 speed up

YOLO v8 m

56 FPS

YOLO v8 m

143 FPS

YOLO v8 m

91 FPS

YOLO v8 m

57 FPS

TensorRT

Half Precision

Half Precision

TensorRT

x3 speed up

Specifications:

    • YOLO v8 m
    • g5.2xlarge
    • batch=64
    • imgsz=1280x1280

51 of 55

YOLOv8, x3 speed up

# Code to optimize

import ultralytics

model = ultralytics.YOLO("yolov8m")

model.export(format="engine", imgsz=1280, batch=64, half=True)

# Or, use https://github.com/NVIDIA-AI-IOT/torch2trt

52 of 55

Consumer PC GPT

53 of 55

Consumer PC GPT

# Code in Python

from gpt4all import GPT4All

model = GPT4All("orca-mini-3b-gguf2-q4_0.gguf")

output = model.generate("The capital of France is ", max_tokens=3)

print(output)

The GPT4All Chat UI supports models from all newer versions of llama.cpp with GGUF models including the Mistral, LLaMA2, LLaMA, OpenLLaMa, Falcon, MPT, Replit, Starcoder, and Bert architectures.

54 of 55

Nomic Vulkan Benchmarks: Single batch item inference token throughput benchmarks.

55 of 55

Q/A