Towards Deterministic End-to-end Latency for Medical AI Systems in NVIDIA Holoscan
Soham Sinha, Shekhar Dwivedi, Mahdi Azizian
NVIDIA, Santa Clara, CA, USA
ICCPS 2024, 05/16/2024
AI and ML in Medical Devices
Medical Imaging and Procedures, Robotic Surgery
2
Sensing | Inference | Data Processing | Visualization | Actuation
NVIDIA Holoscan Platform
What is Holoscan?
3
AI computing platform for real-time streaming and sensor processing applications
Holoscan SDK
Supported in C++ and Python
4
High-level Workflow:�
Sensing
Processing
Actuation
Processing
…
https://github.com/nvidia-holoscan/holoscan-sdk
Operator
Unit of work in Holoscan SDK
Primary Components in Holoscan SDK
5
Types of Holoscan SDK Operators
Sensing Operators�(V4L2, AJA, etc.)
Accelerated Operators�(uses CUDA - Format Converter, etc.)
Inference Operator�(uses CUDA for ML and AI)
Visualization Operator�(uses Vulkan)
A Typical Holoscan Medical AI Application
Endoscopy Tool Tracking
6
Graphics
Video Stream Replayer or Video Card Input
Format Converter
uint8->float32
LSTM TensorRT Inference
Tool Tracking Postprocessor
Holoviz Visualizer
GPU(s)
CUDA Stream
CUDA Stream
Holoscan Data Flow
Graphical Rendering
https://github.com/nvidia-holoscan/holohub/tree/main/applications/endoscopy_tool_tracking
Challenges in Performance Determinism
Multiple Holoscan AI Applications on the same platform
7
GPU
Holoscan AI Application
Compute
Visualization
Holoscan AI Application
Visualization
Resource Contention
Compute
Compute
Compute
Performance Determinism
Our Approach to System Design and Workload Isolation
8
Solution towards a Deterministic System Design for AI Applications
9
Compute GPU
CPU (m … n)
CPU (y … z)
Graphics GPU
SM
SM
SM
SM
SM
SM
SM
SM
…
…
…
…
…
…
SM
CPU (i … j)
SM
SM
SM
…
…
Holoscan AI
Application 1
Holoscan AI
Application 2
Holoscan AI
Application n
Pinned CPU(s)
Compute
Compute
Pinned CPU(s)
Pinned CPU(s)
Compute
Rendering and Visualization
…
…
Isolated Compute
CUDA MPS ACTIVE_THREAD_PERCENTAGE
10
Holoscan AI
Application 1
Holoscan AI
Application 2
Holoscan AI
Application n
SM
SM
SM
SM
…
…
SM
SM
SM
SM
…
…
SM
SM
SM
SM
…
…
Mem
Mem
Mem
Each Holoscan Application gets its own SM and GPU Memory
Evaluation
11
Evaluation: Metrics
How to quantify performance determinism
12
Video Camera
Format Converter
TensorRT Inference
Postprocessor
Holoviz Visualization
End-to-end Latency
Experimental Setup
13
Multiple Instances of Endoscopy Tool Tracking Application
Baseline performance on a single GPU
14
Maximum End-to-end Latency
Latency Tail
Latency Flatness
Metrics increase with more instances
A powerful GPU cannot solve the problem of lower determinism
Our Solution with Workload Isolation
Separate GPUs for Graphics and Compute | Isolated Compute with CUDA MPS
15
Maximum End-to-end Latency
Latency Tail
Latency Flatness
21-30%
21-47%
17-25%
IMG: Isolated Multi-GPU
MPS: CUDA MPS Partition
Do more GPUs lead to better determinism?
16
Short answer: No
Conclusions
17
Future Work
18
Questions?
Thank you!
sohams@nvidia.com
19