Development of CUDA-enabled SU2 Code for GPU Accelerated CFD
An ongoing implementation and study of GPU Support in SU2 Code
SEPTEMBER 2024
SU2 CONFERENCE
Introduction
SEPTEMBER 2024
SU2 CONFERENCE
What we’ve done
What we’ve learnt
Ryzen 7 5800H (Mobile) | 8C/16t, 45W, 16 GB DDR4-3200 RAM |
RTX 3060 (Mobile) | 80 W (Cap), 6GB GDDR6 VRAM, x8 PCIe 3.0 |
Before we begin, a few things
Introduction
SEPTEMBER 2024
SU2 CONFERENCE
Analysis of the Solution Path of SU2 Solvers
SU2 CONFERENCE
SEPTEMBER 2024
What is the use of such a technique in CFD?
GPU Programming
SEPTEMBER 2024
SU2 CONFERENCE
SU2 CPU Code
SEPTEMBER 2024
SU2 CONFERENCE
Profiling done through Tracy Profiler
FGMRES - Solution Methodology
SEPTEMBER 2024
SU2 CONFERENCE
Space Integration | 46% |
Time Integration | 38% |
FGMRES - Solution Methodology
SEPTEMBER 2024
SU2 CONFERENCE
This is the flame graph for the execution of the solver. Preconditioning takes up a major chunk of the time followed by the numerous Matrix Vector Products.
Our target for the change should satisfy the following criteria
Our current candidate for treatment will be Matrix Vector Products.
Implementing CUDA Support in SU2 - Matrix Vector Products
SU2 CONFERENCE
SEPTEMBER 2024
Memory and Thread Management
SEPTEMBER 2024
SU2 CONFERENCE
Images taken from CUDA Documentation
Memory and Thread Management
SEPTEMBER 2024
SU2 CONFERENCE
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block 1
...............................
Total Grid and Problem Size
Each thread in a block represents a single point in the domain.
These threads are all put into blocks of 1024 threads each (Max Block Size)
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block 0
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block 1
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block 2
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block N-3
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block N-2
..........
THREAD
0
THREAD
1
THREAD
2
THREAD
3
THREAD
XDim-3
THREAD
XDim-2
THREAD
XDim -1
Block N-1
XDim = 1024/(nVar*nEqn)
N = nPointDomain/xDim
Grid and Block Initialization
Algorithm
T0, I0
T1, I0
T2, I0
T3, I0
T4, I0
T5, I0
T6, I0
T7, I0
T8, I0
T0, I1
T1, I1
T2, I1
T3, I1
T4, I1
T5, I1
T6, I1
T7, I1
T8, I1
T9, I0
T10, I0
T11, I0
T12, I0
T13, I0
T14, I0
T15, I0
T16, I0
T17, I0
T9, I1
T10, I1
T11, I1
T12, I1
T13, I1
T14, I1
T15, I1
T16, I1
T17, I1
SEPTEMBER 2024
SU2 CONFERENCE
temp_var += matrix[matrix_index + (j * nEqn + k)] * vec[vec_index + k]
atomicAdd(&prod[prod_index + j], temp_var)
Yea sure, I totally know whats going on here
SEPTEMBER 2024
SU2 CONFERENCE
Implementation
SEPTEMBER 2024
SU2 CONFERENCE
Results
SU2 CONFERENCE
SEPTEMBER 2024
Results
SEPTEMBER 2024
SU2 CONFERENCE
GPU = 246.477s
CPU = 201.952
Single Core CPU is 18% faster
GPU = 20.72 ms
CPU = 5.478 ms
Single Core CPU is four times faster
Results
SEPTEMBER 2024
SU2 CONFERENCE
So does this mean that there is no feasibility of GPU Acceleration in SU2?
The answer is no, there is a lot of potential.
SEPTEMBER 2024
SU2 CONFERENCE
Performance Analysis and Solution Proposal
SU2 CONFERENCE
SEPTEMBER 2024
MemCpy 1 | Jacobian Matrix | 1.2 GB/s |
MemCpy 2 | Input Vector | 1.2 GB/S |
MemCpy3 | Resultant Vector | 2 GB/S |
Profiling
SEPTEMBER 2024
SU2 CONFERENCE
Solution
SEPTEMBER 2024
SU2 CONFERENCE
SEPTEMBER 2024
SU2 CONFERENCE
New Algorithm
SEPTEMBER 2024
SU2 CONFERENCE
New Algorithm
There are two approaches we can take to port this entire code
Single Kernel
Multiple Kernel
SEPTEMBER 2024
SU2 CONFERENCE
New Algorithm
We propose a combined approach of the previous two methods
Ease of porting
All of these changes and code have been written in a build version dated September 24, 2024 - which has a faster algorithm, more memory transfer control and better readability
SEPTEMBER 2024
SU2 CONFERENCE
So where is this version?
SEPTEMBER 2024
SU2 CONFERENCE
It doesn’t compile...
SEPTEMBER 2024
SU2 CONFERENCE
Summary
SEPTEMBER 2024
SU2 CONFERENCE
See you (hopefully) for the next SU2 Conference