1 of 13

Towards Precision-Aware Fault ToleranceApproaches for Mixed-Precision Applications

Bo Fang, Siva Kumar Sastry Hari, Timothy Tsai, Xinyi Li, Ganesh Gopalakrishnan,

Ignacio Laguna, Kevin Barker, Ang Li

2 of 13

Mixed-precision Floating Point Computation

  • A trend on today’s emerging applications
    • Tradition scientific domain: 64-bit accuracy
    • AI-driven: 32-bit, 16-bit and even less …

  • Benefits of the reduced precision computation
    • Less memory
    • Faster data transfer
    • Larger arithmetic bandwidth

*

SC22 | Dallas, TX | hpc accelerates.

2

3 of 13

MxP-enabled GEMM Accelerators

*

SC22 | Dallas, TX | hpc accelerates.

3

4 of 13

Non-uniform Resilience Characteristics

*

SC22 | Dallas, TX | hpc accelerates.

4

  • FP values are determined by three components
  • Different formats contain different component-based characteristics

5 of 13

Faults Affecting Bits in Floating Point Values Lead to Different Outcomes

*

SC22 | Dallas, TX | hpc accelerates.

5

[Li et. al SC2017]

[Santos et al. HPCA2019]

High-order exponent bits if corrupted, lead to silent data corruption.

Half precision has less SDC FIT

6 of 13

Goals

  • Build a general evaluation platform for error resilience characterization on MxP-enabled applications

  • Understand the impact of a hardware fault on the error resilience of the application using a particular FP format

  • Design and implement the precision-aware fault tolerance techniques for the emerging applications

*

SC22 | Dallas, TX | hpc accelerates.

6

7 of 13

Extend NVBitFI for Tensor-Core

*

SC22 | Dallas, TX | hpc accelerates.

7

  • A sequence of registers participate
    • Data layout
    • Computation pattern

8 of 13

Data Layout Loaded by Tensor-Core

*

SC22 | Dallas, TX | hpc accelerates.

8

9 of 13

Computation Pattern for Each Thread

*

SC22 | Dallas, TX | hpc accelerates.

9

10 of 13

Fault Injection Methodology

*

SC22 | Dallas, TX | hpc accelerates.

10

 

BF16, FP16 or TF32 multiplication

 

11 of 13

Experimental Setup

  • NVIDIA/cutlass – gemm with different floating-point formats
  • Random v.s. Dedicated fault injection on different FP bits
  • Run 1,000 trials for each configuration
  • NVIDIA V100 and A100 Tensor Cores

*

SC22 | Dallas, TX | hpc accelerates.

11

12 of 13

Preliminary Results

*

SC22 | Dallas, TX | hpc accelerates.

12

13 of 13

Ongoing Work

  • Evaluate the error resilience of machine learning models using different FP formats
    • Effect of 0s
    • Model-specific knowledge
      • CNN v.s. Transformer

*

SC22 | Dallas, TX | hpc accelerates.

13