1 of 32

Fast Machine Learning for Science Workshop 2023

09/26/2023

Ryan Forelli

Real-Time Instability Tracking with Deep Learning on FPGAs in Magnetic Confinement Fusion Devices

2 of 32

  • Fusion: Process of merging atomic nuclei at high temperature/pressure, causing a release of energy while producing minimal waste; immensely promising source of power, but entails significant technical challenges; not commercially viable

Goal: Begin solving the technical challenges facing fusion energy with fast machine learning on hardware.

Overview

2

3 of 32

  • Tokamak: Leading plasma magnetic confinement concept for future fusion reactor designs
    • Vacuum chamber, toroidal in shape
    • Deuterium and tritium plasma injected and heated to millions of degrees Celsius
    • Confinement maintained under generated helical magnetic field
    • Fusion reaction is achieved at sufficiently high temperatures

Magnetic confinement fusion devices: the tokamak

3

4 of 32

The Problem: plasma instability in tokamak devices

  • Magneto-hydrodynamic (MHD) instabilities form when magnetic field lines become distorted and non axi-symmetric and become critical on the order of microseconds
    • Leads to confinement loss on contact with vacuum chamber wall and damage to the reactor
    • One of the major roadblocks preventing lasting sustained fusion reactions

4

5 of 32

The Solution: Active suppression of MHD instabilities

  • Enable accurate real-time tracking with machine learning and non-invasive diagnostics such as high-speed optical imaging
  • Trained CNN to predict sine and cosine components of instability mode
  • Compute new control requests for Tokamak magnetic actuators to suppress the instability in under 20us

5

Tokamak

CNN

[

]

[

]

Controller

DAC

control outputs

actuators

control request

6 of 32

High-speed Camera Setup at the Tokamak

  • Phantom S710 high-speed camera
    • 128x32 resolution
    • 8 or 12-bit depth
    • >50,000fps
    • Note: Currently only camera 1 used
  • Euresys frame grabber
    • CoaXPress protocol
    • 5,000 MB/s camera bandwidth
    • CustomLogic: Your own FPGA logic!

High-speed Camera view

6

7 of 32

CNN training results

  • Input: 32x32 12-bit monochrome image of plasma cross-section
  • Output: predicted sine and cosine component of n=1 mode instability (label (or ground truth) measured from magnetic diagnostics))
  • Selected architecture: convolutional layer block followed by fully-connected dense block
    • Stacking frames and adding second camera yields accuracy improvement

7

n=1 : one toroidal wave period

8 of 32

How do CNNs compare to previous/alternative methods?

  • Previous studies use 1D/2D magnetic data and plasma parameters to track instabilities
  • Alternative methods tested include linear regression and singular value decomposition
  • CNN accuracy improved significantly over previous/alternative methods
    • Input is orders of magnitude larger than previous studies (how will this impact latency?)

8

9 of 32

Hardware deployment

  • Requirements:
    • Inference latency <20us
    • Deterministic timing
    • Minimal data transfer/overhead
    • Power efficient
  • Common deployment solutions: GPUs, CPUs, microcontrollers

control request

Controller

DAC

control outputs

actuators

[

]

[

]

CNN

?

9

10 of 32

  • What about a GPU?
    • Inference latency too high (O(100us)) ❌
    • Non-deterministic timing ❌
    • At least two PCIe hops (O(1ms)) ❌
      • Frame grabber → CPU → GPU
    • Low power efficiency ❌

control request

Controller

DAC

control outputs

actuators

[

]

[

]

CNN

GPU

10

Hardware deployment

Frame grabber

PCIe hop 1

PCIe hop 2

11 of 32

  • Wait… many frame grabbers are FPGA-based!
    • Inference latency satisfactory (<20us) ✔️
    • Deterministic timing ✔️
    • Zero PCIe hops, all computation done on-chip ✔️
    • Highly power efficient ✔️

Coaxial

control request

(RS422)

Controller

DAC

control outputs

host PC

actuators

[

]

[

]

CNN

Euresys CoaxLink Octo

Frame grabber

11

Hardware deployment

12 of 32

  • Frame grabbers typically implement a high speed streaming protocol on an FPGA
    • >70% resources are unused, leaving room for users to implement other image processing functions (or neural networks!)
  • Few manufacturers enable access to firmware
    • Euresys provides a (mostly encrypted) reference design for their frame grabber firmware, enabling access to image data stream
  • Integrate neural network into frame grabber firmware reference design

Framegrabber Model Integration: CustomLogic

CustomLogic

12

13 of 32

Design challenges and constraints

  • Design challenges and constraints imposed by project and hardware
    • Small chip size: Neural network + camera protocol must not exceed chip capacity
    • High clock rate: Design must be routable at 250MHz clock constraints (propagation delays must not exceed clock period); difficult with larger/congested designs
    • Low latency requirement: Delay from frame trigger to exposure to readout to inference to request writeout completion must be <20us

trigger

exposure

frame 0

readout

inference

writeout

t = 0us

t = 20us

Combinational

Logic

In

4ns

clk

xcku035-fbva676-2-e

14 of 32

Additional challenges and inefficiencies

  • Minimum camera resolution limited to 128x32
    • Model input is 32x32
    • Incurs extra streaming latency: ((128*32 - 32*32 - 48) * 4ns) / 8 = 1.5us
  • Camera sensor readout is unnaturally ordered

0

1

2

3

4

5

6

7

0

1

2

3

4

5

6

7

Expected image

Received image

15 of 32

  • Plethora of hardware optimizations and fine-tuning necessary to meet constraints and achieve careful balance between resources and latency

Hardware Optimizations

Optimization

Beneficiary

Quantization aware training (32b → 7b)

Resources

Pruning

Resources

Strategy optimized by layer

Resources, Latency

Reuse factor optimized by layer (see right)

Resources, Latency

Dataflow pipelining (see below)

Latency

Batched input streaming

Latency

Target period, clock unc., & Vivado version scanned

Timing

ReLUs merged

Latency

Synthesis & routing strategies scanned

Timing

Physical optimization looping

Timing

15

16 of 32

Implementation results

  • Model Latency: 7.7us
    • Marked by AXI protocol signals (FILO)
  • Full latency: 17.6us
    • Full trigger → writeout completion latency
  • Resource consumption highlights
    • Highest LUT usage: Conv
    • Highest BRAM Usage: Transport FIFOs
    • Highest DSP usage: Dense
  • Tested at 100kfps

100kfps

8.1us

17 of 32

Pipelining exposure/readout and inference

17

Exposure/readout pipelined with inference

18 of 32

Design verification|firmware

  • How do we verify hardware matches software?
  • Embed model predictions outside region of interest
  • Enables scalable and easy firmware verification with simple python script

128

32

CNN

sine[10:0]

cosine[10:0]

Host PC

32

32

18

19 of 32

Final results and future work

Project summary

  • Completed first implementation toward enabling active control of MHD instability in magnetic confinement fusion devices
    • Model Latency: 7.7us
    • Full latency: 17.6us
    • Tested at 100kfps
  • Broader potential of this work to enable real-time ML inference in any high-speed imaging application using off the shelf frame grabbers

Future Work

  • Integrate current implementation with downstream actuator control system and Tokamak device

19

20 of 32

Acknowledgements

Columbia University - HBT-EP Group

  • Yumou Wei
  • Jeffrey Levesque
  • Chris Hansen

Fermilab - AI/RTPS Division

  • Ryan Forelli
  • Nhan Tran
  • Giuseppe Di Guglielmo
  • Jovan Mitrevski
  • Christian Herwig
  • Jonathan Eisch

20

Supported by US DOE Grant DE-FG02-86ER53222, DE-SC0022234, and US DOE ASCR under the “Real-time Data Reduction Codesign at the Extreme Edge for Science” Project (DE-FOA-0002501).

21 of 32

Appendix A

DAC Input Timing Characteristics

22 of 32

Appendix B

23 of 32

Appendix C

24 of 32

Appendix D

25 of 32

Appendix E

26 of 32

  • Largest layer input consumes model latency
  • Parallelizing the firmware at an IO/streaming level yields additional latency improvements
    • E.g. split input into subimages, passed through two parallel convolutional layers
    • Latency Improvement:

2455 cycles (9.8us) → 1700 cycles (6.8us)

17

17

32

32

conv_0_0

conv_0_1

relu_0_0

relu_0_1

Rearrange

Split

450:450

to

480:420

pool_0_0

pool_0_1

32x17x1

Recombine

30x15x16

30x15x16

30x15x16

30x15x16

32x17x1

30x16x16

30x14x16

15x15x16

15x8x16

15x7x16

+---------------------------+------------------------+------+------+------+------+----------+

| | | Latency | Interval | Pipeline |

| Instance | Module | min | max | min | max | Type |

+---------------------------+------------------------+------+------+------+------+----------+

|conv_2d_cl_1_U0 |conv_2d_cl_1 | 434| 434| 434| 434| none |

|conv_2d_cl_2_U0 |conv_2d_cl_2 | 909| 909| 909| 909| none |

|dense_2_U0 |dense_2 | 61| 62| 61| 62| none |

|relu_U0 |relu | 2| 2| 1| 1| function |

|dense_U0 |dense | 103| 103| 103| 103| none |

|dense_1_U0 |dense_1 | 88| 89| 88| 89| none |

|conv_2d_cl_U0 |conv_2d_cl | 1034| 1034| 1034| 1034| none |

|relu_1_U0 |relu_1 | 2| 2| 1| 1| function |

|relu_2_U0 |relu_2 | 20| 20| 20| 20| none |

|pooling2d_avg_cl_par_7_U0 |pooling2d_avg_cl_par_7 | 521| 521| 521| 521| none |

|pooling2d_cl_U0 |pooling2d_cl | 21| 21| 21| 21| none |

|relu_4_U0 |relu_4 | 904| 904| 904| 904| none |

|relu_3_U0 |relu_3 | 173| 173| 173| 173| none |

|pooling2d_cl_2_U0 |pooling2d_cl_2 | 905| 905| 905| 905| none |

|pooling2d_cl_1_U0 |pooling2d_cl_1 | 174| 174| 174| 174| none |

|myproject_Loop_INPUT_U0 |myproject_Loop_INPUT | 258| 258| 258| 258| none |

|myproject_Loop_POOL_U0 |myproject_Loop_POOL_s | 1026| 1026| 1026| 1026| none |

+---------------------------+------------------------+------+------+------+------+----------+

Appendix F: Extra parallelization: streaming

26

27 of 32

conv2d_0

conv2d_1

conv2d_2

dense_0

dense_1

dense_2

poolings

activations

other (streaming, etc.)

Camera protocol

Appendix G: Logic placement

28 of 32

8.1us

1.6us

Model Serial Readout

sclk

cs

Appendix H: Oscilloscope traces

  • Full serial writeout delay: 1.6us
  • Serial clock and chip-select also implemented

28

old

29 of 32

Appendix I: Camera Readout Fix

  • Reorder pixel stream at model input
    • Model waits longer for first sample
    • Adds 52 cycles of latency

29

Raw readout

Naturally ordered image

CNN

30 of 32

  • VideoIn & VideoIn provide access to data stream & position information (end-of-frame, end-of-line, etc.)
  • Firmware supports ap_uint<8> (8-bit mono)
    • Modified to support ap_uint<12> (12-bit mono)
  • Access to metadata interface provided
  • Computed control requests packed into request_out

CustomLogic Model Integration

void myproject_axi(hls::stream<video_if_t> &VideoIn, hls::stream<video_if_t> &VideoOut,� Metadata_t* MetaIn, Metadata_t* MetaOut, ctrl_req_pack_t &Request) {�� #pragma HLS INTERFACE ap_vld register port=Request�� // Set proper interface for CustomLogic#pragma HLS INTERFACE ap_none port=MetaOut#pragma HLS INTERFACE ap_vld register port=MetaIn// Meta Data valid signal needs to be connect on VideoIn AXI-Stream valid signal#pragma HLS INTERFACE axis depth=0 port=VideoIn#pragma HLS INTERFACE axis port=VideoOut�� (*MetaOut) = meta_data_proc((*MetaIn));�� hls::stream<ctrl_req_pack_t> request_out("request_out");� #pragma HLS STREAM variable=request_out depth=1�� // Call neural network� myproject(VideoIn, VideoOut, request_out);�� // Read control requestctrl_req_pack_t request_v = request_out.read();� Request = request_v;�}

VideoOut

VideoIn

MetaIn

Control Request

MetaOut

HLS top-level function

typedef ap_uint<60> ctrl_req_pack_t;

typedef struct video_struct{� DataMono12 Data;� ap_uint<4> User;�} video_if_t;

Control Request Pack

Data Stream Struct

typedef nnet::array<ap_fixed<11,1>, 2*1> layer21_t;

Prediction Result

31 of 32

Generating a control request

  • Control request computed in firmware
  • Mapped to unsigned binary DAC input
  • Converted to bipolar voltage range (planned)

Coaxial

Pos1 request

[

]

[

]

CNN

Pos2 request

Pos3 request

Pos4 request

Pos5 request

IO Extension Module

op amp

op amp

op amp

op amp

op amp

DIN

V

Control Request

SCLK

CS

Control Request (uint_12)

31

32 of 32

  • C testbench log: cropped/normalized/reordered input, predictions, and control requests
  • RTL simulation includes camera protocol and CustomLogic

Design verification|C/RTL testbench

Cropped Input Frame:

195 195 195 195 195 195 195 195 195 195 ...

Normalized Model Input:

0.764648 0.764648 0.764648 0.764648 0.764648 0.764648 ...

Normalized Reordered Model Input:

0.638672 0.638672 0.638672 0.638672 0.638672 0.638672 ...

Model Predictions:

-0.291992 0.140625

Control Request:

coil1 = 0.481445 | coil2 = 0.416016 | coil3 = 0.1875 | coil4 = -0.110352 | coil5 = -0.368164 |

Control Request Mapped to DAC (0-4095):

coil1 = 3034 = 101111011010 | coil2 = 2900 = 101101010100 | coil3 = 2432 = 100110000000 | coil4 = 1822 = 011100011110 | coil5 = 1294 = 010100001110 |

Request Unsigned Binaary Packed

010100001110 011100011110 100110000000 101101010100 101111011010

Received Image:

192 192 192 192 192 192 192 192 192 192 192 192 ...

METADATA:

StreamId = 0x00

SourceTag = 0x0000

Xsize = 0x000080

Xoffs = 0x000000

Ysize = 0x000020

Yoffs = 0x000000

DsizeL = 0x000020

PixelF = 0x0101

TapG = 0x0000

Flags = 0x00

Timestamp = 0x00000001

PixProcFlgs = 0x00

Status = 0x00000000

PAYLOAD:

00 01 02 03 04 05 06 07 08 09 0A 0B 0C 0D 0E 0F 10 ...

32