Towards Precision-Aware Fault Tolerance�Approaches for Mixed-Precision Applications
Bo Fang, Siva Kumar Sastry Hari, Timothy Tsai, Xinyi Li, Ganesh Gopalakrishnan,
Ignacio Laguna, Kevin Barker, Ang Li
Mixed-precision Floating Point Computation
*
SC22 | Dallas, TX | hpc accelerates.
2
MxP-enabled GEMM Accelerators
*
SC22 | Dallas, TX | hpc accelerates.
3
Non-uniform Resilience Characteristics
*
SC22 | Dallas, TX | hpc accelerates.
4
Faults Affecting Bits in Floating Point Values Lead to Different Outcomes
*
SC22 | Dallas, TX | hpc accelerates.
5
[Li et. al SC2017]
[Santos et al. HPCA2019]
High-order exponent bits if corrupted, lead to silent data corruption.
Half precision has less SDC FIT
Goals
*
SC22 | Dallas, TX | hpc accelerates.
6
Extend NVBitFI for Tensor-Core
*
SC22 | Dallas, TX | hpc accelerates.
7
Data Layout Loaded by Tensor-Core
*
SC22 | Dallas, TX | hpc accelerates.
8
Computation Pattern for Each Thread
*
SC22 | Dallas, TX | hpc accelerates.
9
Fault Injection Methodology
*
SC22 | Dallas, TX | hpc accelerates.
10
BF16, FP16 or TF32 multiplication
Experimental Setup
*
SC22 | Dallas, TX | hpc accelerates.
11
Preliminary Results
*
SC22 | Dallas, TX | hpc accelerates.
12
Ongoing Work
*
SC22 | Dallas, TX | hpc accelerates.
13