1 of 19

Towards Deterministic End-to-end Latency for Medical AI Systems in NVIDIA Holoscan

Soham Sinha, Shekhar Dwivedi, Mahdi Azizian

NVIDIA, Santa Clara, CA, USA

ICCPS 2024, 05/16/2024

2 of 19

AI and ML in Medical Devices

Medical Imaging and Procedures, Robotic Surgery

2

Sensing | Inference | Data Processing | Visualization | Actuation

3 of 19

NVIDIA Holoscan Platform

What is Holoscan?

3

AI computing platform for real-time streaming and sensor processing applications

4 of 19

Holoscan SDK

Supported in C++ and Python

4

High-level Workflow:

  • Acquire streaming data from sensors�
  • Process data in a pipeline of tasks�
  • Display or output based on processed data

Sensing

Processing

Actuation

Processing

https://github.com/nvidia-holoscan/holoscan-sdk

Operator

Unit of work in Holoscan SDK

5 of 19

Primary Components in Holoscan SDK

5

Types of Holoscan SDK Operators

Sensing Operators�(V4L2, AJA, etc.)

Accelerated Operators�(uses CUDA - Format Converter, etc.)

Inference Operator�(uses CUDA for ML and AI)

Visualization Operator�(uses Vulkan)

6 of 19

A Typical Holoscan Medical AI Application

Endoscopy Tool Tracking

6

Graphics

Video Stream Replayer or Video Card Input

Format Converter

uint8->float32

LSTM TensorRT Inference

Tool Tracking Postprocessor

Holoviz Visualizer

GPU(s)

CUDA Stream

CUDA Stream

Holoscan Data Flow

Graphical Rendering

https://github.com/nvidia-holoscan/holohub/tree/main/applications/endoscopy_tool_tracking

7 of 19

Challenges in Performance Determinism

Multiple Holoscan AI Applications on the same platform

7

GPU

Holoscan AI Application

Compute

Visualization

Holoscan AI Application

Visualization

Resource Contention

Compute

Compute

Compute

Performance Determinism

  • Concurrent CUDA and Graphics Workloads across multiple Holoscan application processes
    • suffers from resource contention for multiple AI and ML workloads sharing GPU compute resources
    • context switch overhead between compute and graphics workloads
    • performance determinism degrades

8 of 19

Our Approach to System Design and Workload Isolation

8

9 of 19

Solution towards a Deterministic System Design for AI Applications

9

Compute GPU

CPU (m … n)

CPU (y … z)

Graphics GPU

SM

SM

SM

SM

SM

SM

SM

SM

SM

CPU (i … j)

SM

SM

SM

Holoscan AI

Application 1

Holoscan AI

Application 2

Holoscan AI

Application n

Pinned CPU(s)

Compute

Compute

Pinned CPU(s)

Pinned CPU(s)

Compute

Rendering and Visualization

  • Dedicated CPUs per application�
  • Dedicated GPU SMs for ML and AI Compute per application�
  • Isolated Compute and Graphics

10 of 19

Isolated Compute

CUDA MPS ACTIVE_THREAD_PERCENTAGE

10

Holoscan AI

Application 1

Holoscan AI

Application 2

Holoscan AI

Application n

SM

SM

SM

SM

SM

SM

SM

SM

SM

SM

SM

SM

Mem

Mem

Mem

Each Holoscan Application gets its own SM and GPU Memory

11 of 19

Evaluation

11

12 of 19

Evaluation: Metrics

How to quantify performance determinism

12

  • End-to-end Latency
    • the time between the receival of a message at the input and publication of a message at the output
    • Average end-to-end Latency
      • Standard deviation
    • Maximum (worst-case) end-to-end latency
    • Median and Minimum��
  • Latency Tail
    • Difference between 95 and 100 percentile latency�
  • Latency Flatness
    • Difference between 10 and 90 percentile latency
    • Concentration of observed latencies

Video Camera

Format Converter

TensorRT Inference

Postprocessor

Holoviz Visualization

End-to-end Latency

13 of 19

Experimental Setup

  • Medical AI Applications
    • Endoscopy Tool Tracking: tool identification and tracking
    • Ultrasound Segmentation: segmentation of the spine
    • Multi-AI Ultrasound: critical linear measurements of the heart, cardiac anatomy classification, etc.�
  • Hardware
    • x86 Desktop Workstation
    • GPUs
      • 1x RTX A4000
      • 1x RTX A6000
      • 2x RTX A4000
      • 1x RTX A4000 + 1x RTX A6000

13

14 of 19

Multiple Instances of Endoscopy Tool Tracking Application

Baseline performance on a single GPU

14

Maximum End-to-end Latency

Latency Tail

Latency Flatness

Metrics increase with more instances

A powerful GPU cannot solve the problem of lower determinism

15 of 19

Our Solution with Workload Isolation

Separate GPUs for Graphics and Compute | Isolated Compute with CUDA MPS

15

Maximum End-to-end Latency

Latency Tail

Latency Flatness

21-30%

21-47%

17-25%

IMG: Isolated Multi-GPU

MPS: CUDA MPS Partition

16 of 19

Do more GPUs lead to better determinism?

16

  • 2x A4000 (C+G): Compute and Graphics both are on �two RTX A4000 GPUs�
  • IMG-MPS (2x A4000): 1x A4000 for Compute and 1x A4000�for Graphics + MPS on Compute GPU�
  • Better Workload Isolation design leads to
    • 35% decrement in maximum end-to-end latency
    • 42% efficient GPU utilization

Short answer: No

17 of 19

Conclusions

  • Medical AI applications demand easy usage of ML, visualization, sensing and actuation capabilities�
  • NVIDIA Holoscan platform provides these capabilities off-the-shelf in a production-ready environment�
  • When multiple such medical AI applications are running, performance determinism takes a hit�
  • Using CUDA MPS and isolated GPUs for compute and graphics provides better end-to-end latency predictability

17

18 of 19

Future Work

18

  • Incorporate similar benefits on IGX Orin embedded platform
    • Isolated Compute
    • Isolated Compute and Graphics�
  • Hardware virtualization capabilities for Isolation�
  • Better determinism on a single GPU for mixed compute and graphics workloads

19 of 19

Questions?

Thank you!

sohams@nvidia.com

19