1
Petros Toupas, Zhewen Yu�Supervisor: Prof. Christos-Savvas Bouganis
SMOF: Streaming Modern CNNs on FPGAs with
Smart Off-Chip Eviction
intelligent Digital Systems Lab
Dept. Of Electrical and Electronic Engineering
www.imperial.ac.uk/idsl
intelligent Digital Systems Lab
2
Deep Leaning & Computer Vision
Classification
Object Detection
Image Segmentation
Action Recognition
intelligent Digital Systems Lab
3
ML & Computer Vision Applications
Automotive
Robotics
Manufacturing
Healthcare
Security
intelligent Digital Systems Lab
4
Applications Requirements
On the Edge
On Datacenters
Low Latency
Low Power
High Throughput
Low Power
FPGAs
Customizability
Power Efficiency
Adaptability
intelligent Digital Systems Lab
5
AI Hardware Accelerators
Single Computation Engines:
Streaming Architectures:
Images source: Stylianos I. Venieris, Alexandros Kouris, and Christos-Savvas Bouganis. 2018. Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions. ACM Comput. Surv. 51, 3, Article 56 (May 2019)
intelligent Digital Systems Lab
6
Streaming Architectures Challenges & Opportunities
Part 1
Part 2
Part 3
Model
FPGA �reconfiguration
FPGA �reconfiguration
intelligent Digital Systems Lab
7
Deep Skip Connections on modern CNNs:
Deep Skip Connections
(7 to 40+)
UNet (segmentation):
Deep Skip Connections
(20 to 60+)
YOLO�(object detection):
Deep Skip Connections
(5 to 10+)
X3D�(action recognition):
Streaming Architectures Challenge #1
intelligent Digital Systems Lab
8
Large number of Parameters on modern CNNs:
Streaming Architectures Challenge #2
UNet Memory Breakdown | ||
Device | U200 | |
Total BRAM | 4320 | |
Total URAM | 960 | |
BRAM (weights) | 2854 (66.1%) | |
BRAM (skip connections) | 249 (5.7%) | |
BRAM (sliding windows) | 551 (12.7%) | |
Total BRAM (used) | 3654 (84.6%) | |
URAM (weights) | 864 (90%) | |
Total URAM (used) | 864 (90%) | |
Model | Parameters (M) | MB (FP32) |
Yolov8n | 3.16 | 12.05 |
UNet | 28.96 | 110.47 |
X3D-M | 3.82 | 14.57 |
UNET3D | 5.65 | 21.55 |
intelligent Digital Systems Lab
9
Contribution 1: Activation Eviction & Weight Fragmentation
Buffers necessary on branch connections & sliding windows in Conv layers:
Explore the trade-off between on-chip FIFOs and off-chip Bandwidth
Weights fragmented into static and dynamic regions:
Explore the ratio between static and dynamic regions of weights memory on Conv layers
Arbitrary activation eviction points:
Arbitrary weights offloading to off-chip:
intelligent Digital Systems Lab
10
Contribution 2: Subgraph-based partitioning methodology
intelligent Digital Systems Lab
11
Outcomes
Work | Snowflake | Angel-eye | Brainwave | Vitis AI | DeepBurning | FINN | DeepBuilder | HPIPE | SMOF (Ours) |
Architecture Style | Single Engine Architectures | Streaming Architectures | |||||||
Classification | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Detection | 🗴 | ✓ | 🗴 | ✓ | 🗴 | 🗴 | ✓ | ✓ | ✓ |
Segmentation | 🗴 | 🗴 | 🗴 | ✓ | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
Action Recognition | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
2D CNN | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
3D CNN | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
intelligent Digital Systems Lab
12
UNet Accelerator Design
Latency | Invalid Design |
Throughput | Invalid Design |
DSP | 6062 (89%) |
BRAM | 4580 (106.5%) |
URAM | 928 (96.67%) |
LUT | 1018.8K (86.2%) |
FF | 841K (36%) |
BW | 1.2 Gbps (0.2%) |
intelligent Digital Systems Lab
13
UNet Accelerator Design
Latency | 47 ms |
Throughput | 21 fps |
DSP | 6062 (89%) |
BRAM | 3654 (85%) |
URAM | 864 (90%) |
LUT | 1040K (88%) |
FF | 841K (36%) |
BW | 225 Gbps (37%) |
intelligent Digital Systems Lab
14
DSE
Automated Design Space Exploration (DSE)
Objective:
Latency or Throughput
Constraints:
Resources (both on and off chip)
intelligent Digital Systems Lab
15
DSE Variables
Decision required for each layer in the graph!
Partitioning into Subgraphs
Activation eviction & �Weight fragmentation
Layer Parallelism
New Feature
New Feature
Feature
intelligent Digital Systems Lab
16
DSE Optimisation
Initialise Resource Minimal Design
Parallelism Allocation
Compute (DSP) Constraints
Reset to Previous Valid Design
On-Chip Memory�Constraints
Off-Chip Bandwidth�Constraints
Re-Allocate Memory
Y
N
Y
Y
N
N
intelligent Digital Systems Lab
17
Ablation Study (UNet Model)
Baseline
Off-Chip Streaming
Off-Chip Streaming Compression
intelligent Digital Systems Lab
18
Evaluated Models Characteristics
2D CNNs
3D CNNs
intelligent Digital Systems Lab
19
Evaluation (Comparison with SoA)
10.65 x
intelligent Digital Systems Lab
20
Evaluation (Comparison with SoA)
No improvement:
intelligent Digital Systems Lab
21
Evaluation (Comparison with SoA)
2 x
2.8 x
intelligent Digital Systems Lab
22
Evaluation (Comparison with SoA)
GOP/s -> 1.7 x
GOP/s/DSP -> No improvement
DSP implementation variations between Intel and AMD FPGAs
intelligent Digital Systems Lab
23
Evaluation (Comparison with SoA)
FPS -> 23 x
GOP/s/DSP -> 0.89 x
intelligent Digital Systems Lab
24
Evaluation (Comparison with SoA)
GOP/s -> 1.98 x
First Time Ever Mapped
intelligent Digital Systems Lab
25
Conclusion
Work | Snowflake | Angel-eye | Brainwave | Vitis AI | DeepBurning | FINN | DeepBuilder | HPIPE | SMOF (Ours) |
Architecture Style | Single Engine Architectures | Streaming Architectures | |||||||
Classification | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Detection | 🗴 | ✓ | 🗴 | ✓ | 🗴 | 🗴 | ✓ | ✓ | ✓ |
Segmentation | 🗴 | 🗴 | 🗴 | ✓ | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
Action Recognition | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
2D CNN | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
3D CNN | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | 🗴 | ✓ |
SMOF overcomes streaming architectures limitations:
intelligent Digital Systems Lab
Petros Toupas
iDSL website
intelligent Digital Systems Lab