1 of 26

1

Petros Toupas, Zhewen Yu�Supervisor: Prof. Christos-Savvas Bouganis

SMOF: Streaming Modern CNNs on FPGAs with

Smart Off-Chip Eviction

intelligent Digital Systems Lab

Dept. Of Electrical and Electronic Engineering

www.imperial.ac.uk/idsl

intelligent Digital Systems Lab

2 of 26

2

Deep Leaning & Computer Vision

Classification

Object Detection

Image Segmentation

Action Recognition

intelligent Digital Systems Lab

3 of 26

3

ML & Computer Vision Applications

Automotive

  • Pedestrian Detection
  • Collision Avoidance

Robotics

  • Autonomous Navigation
  • Human Robot�Interaction (HRI)

Manufacturing

  • Product line inspection�& assessment
  • Detection of defective�objects

Healthcare

  • Elderly Activity Monitoring
  • Rehabilitation Support

Security

  • Surveillance Systems
  • Anomaly Detection

intelligent Digital Systems Lab

4 of 26

4

Applications Requirements

On the Edge

On Datacenters

Low Latency

Low Power

High Throughput

Low Power

FPGAs

Customizability

Power Efficiency

Adaptability

intelligent Digital Systems Lab

5 of 26

5

AI Hardware Accelerators

Single Computation Engines:

    • Systolic Array of Processing Elements (PEs)
    • Time shared between layers
    • Memory Bounded (most of the times)
    • Angel-Eye, Snowflake, FP-DNN

Streaming Architectures:

    • Data driven pipeline execution
    • Hardware tailored to each layer
    • Compute Bounded (most of the times)
    • fpgaConvNet, FINN, HPIPE

Images source: Stylianos I. Venieris, Alexandros Kouris, and Christos-Savvas Bouganis. 2018. Toolflows for Mapping Convolutional Neural Networks on FPGAs: A Survey and Future Directions. ACM Comput. Surv. 51, 3, Article 56 (May 2019)

intelligent Digital Systems Lab

6 of 26

6

Streaming Architectures Challenges & Opportunities

  • Limited by on-chip memory �(weights + buffering)

  • Mapping modern (large) CNNs requires:
    • Partitioning of the model
    • Device reconfiguration�
  • Limiting the coverage of application requirements (e.g., low latency)

  • Underutilized off-chip memory and bandwidth

Part 1

Part 2

Part 3

Model

FPGA �reconfiguration

FPGA �reconfiguration

intelligent Digital Systems Lab

7 of 26

7

Deep Skip Connections on modern CNNs:

    • Significant on-chip storage for buffering

Deep Skip Connections

(7 to 40+)

UNet (segmentation):

Deep Skip Connections

(20 to 60+)

YOLO�(object detection):

Deep Skip Connections

(5 to 10+)

X3D�(action recognition):

Streaming Architectures Challenge #1

intelligent Digital Systems Lab

8 of 26

8

Large number of Parameters on modern CNNs:

    • Significant on-chip storage for weights and buffering activations

Streaming Architectures Challenge #2

UNet Memory Breakdown

Device

U200

Total BRAM

4320

Total URAM

960

BRAM (weights)

2854 (66.1%)

BRAM (skip connections)

249 (5.7%)

BRAM (sliding windows)

551 (12.7%)

Total BRAM (used)

3654 (84.6%)

URAM (weights)

864 (90%)

Total URAM (used)

864 (90%)

Model

Parameters (M)

MB (FP32)

Yolov8n

3.16

12.05

UNet

28.96

110.47

X3D-M

3.82

14.57

UNET3D

5.65

21.55

intelligent Digital Systems Lab

9 of 26

9

Contribution 1: Activation Eviction & Weight Fragmentation

Buffers necessary on branch connections & sliding windows in Conv layers:

    • Balance data flow on branches to prevent stalls
    • Prevent possible deadlocks in execution

Explore the trade-off between on-chip FIFOs and off-chip Bandwidth

Weights fragmented into static and dynamic regions:

  • Static regions are always on-chip and �read-only 
  • Dynamic regions sharing same on-chip space in a time-multiplexed manner

Explore the ratio between static and dynamic regions of weights memory on Conv layers

Arbitrary activation eviction points:

Arbitrary weights offloading to off-chip:

intelligent Digital Systems Lab

10 of 26

10

Contribution 2: Subgraph-based partitioning methodology

  • Arbitrary reconfiguration points

  • Flexible partitioning into subgraphs
    • Multiple off-chip input connections
    • Multiple off-chip output connections

  • Exploring the trade-off between the latency and the throughput of the system during DSE

intelligent Digital Systems Lab

11 of 26

11

Outcomes

  • Addressed the limitations in scaling networks with large parameters and long skip connections

Work

Snowflake

Angel-eye

Brainwave

Vitis AI

DeepBurning

FINN

DeepBuilder

HPIPE

SMOF (Ours)

Architecture Style

Single Engine Architectures

Streaming Architectures

Classification

Detection

🗴

🗴

🗴

🗴

Segmentation

🗴

🗴

🗴

🗴

🗴

🗴

🗴

Action Recognition

🗴

🗴

🗴

🗴

🗴

🗴

🗴

🗴

2D CNN

3D CNN

🗴

🗴

🗴

🗴

🗴

🗴

🗴

🗴

  • Support for bigger CNN model graphs by reducing:
    • On-chip memory storage for weights
    • On-chip memory storage for buffering (long skip connections, sliding windows for Conv)

intelligent Digital Systems Lab

12 of 26

12

UNet Accelerator Design

Latency

Invalid Design

Throughput

Invalid Design

DSP

6062 (89%)

BRAM

4580 (106.5%)

URAM

928 (96.67%)

LUT

1018.8K (86.2%)

FF

841K (36%)

BW

1.2 Gbps (0.2%)

intelligent Digital Systems Lab

13 of 26

13

UNet Accelerator Design

Latency

47 ms

Throughput

21 fps

DSP

6062 (89%)

BRAM

3654 (85%)

URAM

864 (90%)

LUT

1040K (88%)

FF

841K (36%)

BW

225 Gbps (37%)

intelligent Digital Systems Lab

14 of 26

14

DSE

Automated Design Space Exploration (DSE)

Objective:

Latency or Throughput

Constraints:

Resources (both on and off chip)

intelligent Digital Systems Lab

15 of 26

15

DSE Variables

Decision required for each layer in the graph!

Partitioning into Subgraphs

Activation eviction & �Weight fragmentation

Layer Parallelism

New Feature

New Feature

Feature

intelligent Digital Systems Lab

16 of 26

16

DSE Optimisation

Initialise Resource Minimal Design

Parallelism Allocation

Compute (DSP) Constraints

Reset to Previous Valid Design

On-Chip Memory�Constraints

Off-Chip Bandwidth�Constraints

Re-Allocate Memory

Y

N

Y

Y

N

N

intelligent Digital Systems Lab

17 of 26

17

Ablation Study (UNet Model)

Baseline

  • Vanilla fpgaConvNet
  • 8-bit BFP for both weights and activations

Off-Chip Streaming

    • Activations Only (4.4%)
    • Weights Only (22.2%)
    • Activations + Weights (31.1%)

Off-Chip Streaming Compression

    • No encoding
    • Huffman (20.3%)
    • RLE (83.1%)

intelligent Digital Systems Lab

18 of 26

18

Evaluated Models Characteristics

2D CNNs

  • Unet (image segmentation)
  • Yolov8n (object detection)

3D CNNs

  • UNet3D (3D volumetric segmentation)
  • X3D-M (action recognition)

intelligent Digital Systems Lab

19 of 26

19

Evaluation (Comparison with SoA)

10.65 x

intelligent Digital Systems Lab

20 of 26

20

Evaluation (Comparison with SoA)

No improvement:

  • UNet (130 vs 5.6 MACs) -> 23.3x
  • ZCU102 limited resources
  • Compute bounded

intelligent Digital Systems Lab

21 of 26

21

Evaluation (Comparison with SoA)

2 x

2.8 x

intelligent Digital Systems Lab

22 of 26

22

Evaluation (Comparison with SoA)

GOP/s -> 1.7 x

GOP/s/DSP -> No improvement

DSP implementation variations between Intel and AMD FPGAs

intelligent Digital Systems Lab

23 of 26

23

Evaluation (Comparison with SoA)

FPS -> 23 x

GOP/s/DSP -> 0.89 x

intelligent Digital Systems Lab

24 of 26

24

Evaluation (Comparison with SoA)

GOP/s -> 1.98 x

First Time Ever Mapped

intelligent Digital Systems Lab

25 of 26

25

Conclusion

Work

Snowflake

Angel-eye

Brainwave

Vitis AI

DeepBurning

FINN

DeepBuilder

HPIPE

SMOF (Ours)

Architecture Style

Single Engine Architectures

Streaming Architectures

Classification

Detection

🗴

🗴

🗴

🗴

Segmentation

🗴

🗴

🗴

🗴

🗴

🗴

🗴

Action Recognition

🗴

🗴

🗴

🗴

🗴

🗴

🗴

🗴

2D CNN

3D CNN

🗴

🗴

🗴

🗴

🗴

🗴

🗴

🗴

  • Novel memory optimisation methodology utilising both on-chip and off-chip memory
  • Partially offloading weights and activations to off-chip without penalizing the computation pipeline
  • A subgraph partitioning methodology allowing flexible trade-offs between latency and throughput

SMOF overcomes streaming architectures limitations:

intelligent Digital Systems Lab

26 of 26

pt121@ic.ac.uk

Petros Toupas

 iDSL website

intelligent Digital Systems Lab