1 of 14

Fast Machine Learning for Science Conference 2024

Purdue University

10/16/2024

An open platform for in-situ high-speed computer vision with hls4ml

2 of 14

Background and prerequisite work

2

  • Fields that rely on high speed imaging analysis:
    • Material science — 4D STEM structure detection, RHEED, defect dynamics, etc.
    • Biological cell-sorting & microfluidics
    • Manufacturing — QC, early fault diagnosis, etc.
    • Fusion MHD — Instability suppression in Tokamak fusion reactors
      • Deployed CNNs to frame grabber devices for low-latency control - achieved
        • Achieved 17.6us trigger to output, 7.7us model latency

Credit: Dr. Josh Agar

3 of 14

An open platform for in-situ high-speed computer vision

3

  • Objective: Streamline deployment of standard streaming hls4ml neural networks to off-the-shelf camera hardware for high-speed imaging applications
    • Leveraged open reference design from manufacturer: CustomLogic
    • Developed HLS wrappers and necessary auxiliary files to…
      • integrate hls4ml IP into frame grabber design
      • unpack and crop user-defined region of interest from CoaXPress camera stream
        • parse axi-stream side channel info for spatial information
      • reorder camera stream
      • duplicate stream
      • attach predictions to frame
      • serial output via TTL IO
      • enable easy hw benchmarking via frame grabber TTL IO
      • build Vivado project and synthesize design to meet 250MHz timing constraints
    • and more…

hls4ml 4 FG wrapper

hls4ml 4 FG wrapper

© euresys

4 of 14

Tutorials

4

Not one… but TWO tutorials to get you started:

  1. Basic deployment and benchmarking tutorial (Jupyter & Medium format)
    1. The basics, SW/HW requirements
    2. Pretrained quantized MNIST model
    3. Benchmarking
  2. Advanced YOLO & hls4ml tutorial
    • Fake YOLO (FOLO) model (reconfigurable for your application)
    • Dataset generation
    • Quantization-aware training
    • hls4ml hardware optimization API for pattern pruning
    • hls4ml extensions API for custom data reduction layer
    • Post-training quantization
    • hls4ml profiling tool
    • hls4ml FIFO optimization
    • Frame grabber firmware integration (most automated)
    • Benchmarking

5 of 14

PR & hls4ml Demo

5

Already released

  • Medium article
  • Euresys Vision in the Loop case study article

In the works:

  • Card throwing demo with Rick Smith Jr.

6 of 14

6

An open platform for in-situ high-speed computer vision with hls4ml

Existing applications

  • Fusion
  • RHEED gaussian fitting (FOLO)
  • Cell sorting (FOLO)
  • 4D TEM crystal structure detection

Future work

  • Support QSFP boards
  • Card-throwing demo

Conclusion

  • Two comprehensive tutorials
  • Two PR articles

Thanks!

7 of 14

Acknowledgements

7

Supported by US DOE Grant DE-FG02-86ER53222, DE-SC0022234, and US DOE ASCR under the “Real-time Data Reduction Codesign at the Extreme Edge for Science” Project (DE-FOA-0002501).

Ryan F. Forelli1,2,3, Nhan Tran1,3 Josh C. Agar4 Yumou Wei5, Chris Hansen5, Jeffrey Levesque5,

Michael Cyros6 Eric Janssen6 Paulo Possa6 Benoît Trémérie6

1Fermi National Accelerator Laboratory, Batavia, IL

2Lehigh University, Bethlehem, PA

3Northwestern University, Evanston, IL

4Drexel University, Philadelphia, PA

5Columbia University, New York City, NY

6Euresys S.A., Seraing, Belgium

Advisor: Professor Seda Ogrenci

8 of 14

Appendix A

9 of 14

Design verification|firmware

  • How do we verify hardware matches software?
  • Embed model predictions outside region of interest
  • Enables scalable and easy firmware verification with simple python script

128

32

CNN

sine[10:0]

cosine[10:0]

Host PC

32

32

9

10 of 14

Pipelining exposure/readout and inference

10

Exposure/readout pipelined with inference

11 of 14

Additional challenges and inefficiencies

  • Minimum camera resolution limited to 128x32
    • Model input is 32x32
    • Incurs extra streaming latency: ((128*32 - 32*32 - 48) * 4ns) / 8 = 1.5us
  • Camera sensor readout is unnaturally ordered

0

1

2

3

4

5

6

7

0

1

2

3

4

5

6

7

Expected image

Received image

12 of 14

Fusion Project|Constraints & Optimization

12

  • Strict constraints: latency, resources, timing (250MHz)
  • Necessitates hardware optimizations and fine-tuning to achieve careful balance between resources and latency

Optimization

Beneficiary

Quantization aware training (32b → 7b)

Resources

Pruning

Resources

Strategy optimized by layer

Resources, Latency

Reuse factor optimized by layer (see right)

Resources, Latency

Dataflow pipelining (see appendix)

Latency

Batched input streaming

Latency

Target period, clock unc., & Vivado version scanned

Timing

ReLUs merged

Latency

Synthesis & routing strategies scanned

Timing

Physical optimization looping

Timing

Inference latency: 7.7us

Total latency: 17.6us

Throughput: 100kfps

13 of 14

Fusion|Implementation Results

  • Model Latency: 7.7us
    • Includes streaming
  • Full latency: 17.6us

14 of 14

  • Removed input pooling, crop input image instead
  • conv_0 is currently the bottleneck
  • Parallelize conv by streaming overlapping subimages to separate conv2d layers
    • Diminishing returns as parallelization factor increases due to overlap needed for filter
    • To avoid incurring streaming latency (1024 cycles), wait to recombine resulting streams until after pooling layer
  • Stream result into parallel ReLUs
  • Parallelize pooling, apply an uneven split
  • conv_1 & conv_2 reuse factors decreased
  • Latency Improvement: 2455 cycles (9.8us) → 1700 cycles (6.8us)

Fusion|Conv2D Parallelization

17

17

32

32

conv_0_0

conv_0_1

relu_0_0

relu_0_1

Rearrange

Split

450:450

to

480:420

pool_0_0

pool_0_1

32x17x1

Recombine

30x15x16

30x15x16

30x15x16

30x15x16

32x17x1

30x16x16

30x14x16

15x15x16

15x8x16

15x7x16

+---------------------------+------------------------+------+------+------+------+----------+

| | | Latency | Interval | Pipeline |

| Instance | Module | min | max | min | max | Type |

+---------------------------+------------------------+------+------+------+------+----------+

|conv_2d_cl_1_U0 |conv_2d_cl_1 | 434| 434| 434| 434| none |

|conv_2d_cl_2_U0 |conv_2d_cl_2 | 909| 909| 909| 909| none |

|dense_2_U0 |dense_2 | 61| 62| 61| 62| none |

|relu_U0 |relu | 2| 2| 1| 1| function |

|dense_U0 |dense | 103| 103| 103| 103| none |

|dense_1_U0 |dense_1 | 88| 89| 88| 89| none |

|conv_2d_cl_U0 |conv_2d_cl | 1034| 1034| 1034| 1034| none |

|relu_1_U0 |relu_1 | 2| 2| 1| 1| function |

|relu_2_U0 |relu_2 | 20| 20| 20| 20| none |

|pooling2d_avg_cl_par_6_U0 |pooling2d_avg_cl_par_6 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_5_U0 |pooling2d_avg_cl_par_5 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_4_U0 |pooling2d_avg_cl_par_4 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_3_U0 |pooling2d_avg_cl_par_3 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_2_U0 |pooling2d_avg_cl_par_2 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_1_U0 |pooling2d_avg_cl_par_1 | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_U0 |pooling2d_avg_cl_par | 521| 521| 521| 521| none |

|pooling2d_avg_cl_par_7_U0 |pooling2d_avg_cl_par_7 | 521| 521| 521| 521| none |

|pooling2d_cl_U0 |pooling2d_cl | 21| 21| 21| 21| none |

|relu_4_U0 |relu_4 | 904| 904| 904| 904| none |

|relu_3_U0 |relu_3 | 173| 173| 173| 173| none |

|pooling2d_cl_2_U0 |pooling2d_cl_2 | 905| 905| 905| 905| none |

|pooling2d_cl_1_U0 |pooling2d_cl_1 | 174| 174| 174| 174| none |

|myproject_Loop_INPUT_U0 |myproject_Loop_INPUT | 258| 258| 258| 258| none |

|myproject_Loop_POOL_U0 |myproject_Loop_POOL_s | 1026| 1026| 1026| 1026| none |

+---------------------------+------------------------+------+------+------+------+----------+