1 of 12

SIMD 64x2

2 of 12

Benchmarks

  • Microbenchmarks selected from Dart SIMD paper [0]
    • int64_average
    • double_average
    • double_sum
    • 4x4 matrix_multiply
  • Mandelbrot set (from language benchmarks game)
  • Computation in Eigen (a = a + b * c)
  • SIMD-oriented Fast Mersenne Twister (SFMT) 💡

3 of 12

Test environment

  • Linux workstation, x64 (Xeon(R) E5 28 cores)
  • Pixel 3, arm64 (sdm845, Cortex-A75)

4 of 12

Results

x64

arm64

5 of 12

Results

x64

arm64

6 of 12

Cortex-A76 v.s. Cortex-A75 💡

Cortex-A76 - FMUL latency = 3, throughput = 2

Cortex-A75 - FMUL latency = 3, throughput = 1

Cortex-A75 - Scalar FMUL latency = 3 throughput 2

7 of 12

Cortex-A76 v.s. Cortex-A75 💡

Instruction

Latency

Throughput (instr/cycle)

A75 Scalar FMUL

3

2

A75 SIMD FMUL

3

1

A76 Scalar FMUL

3

2

A76 SIMD FMUL

3

2

8 of 12

Distribution of instructions

Instructions

Count

f64x2.mul

193

f64x2.add

188

f64x2.sub

51

i64x2.shl 💡

26

f64x2.splat

24

i64x2.shr_u 💡

23

i64x2.splat

17

9 of 12

Ongoing work

  • Real-world workloads for i64 instructions 💡
    • Arithmetic, comparisons, min/max
  • Real-world workloads for f64 instructions
    • Comparison, min/max
  • some more x64 instructions
    • f64x2.convert.i64x2
    • i64x2.trunc.f64x2
    • f64x2.sqrt
  • arm64 implementation halfway done
  • int64 workloads (blake2b)
    • Xor, add, shr

10 of 12

Future work

  • 32-bit architectures (if and when 64x2 instructions go into spec)
  • https://github.com/WebAssembly/benchmarks ?

11 of 12

Call for contributions

  • Send benchmarks (preferably code)
    • porting existing benchmarks (from xmmintrin.h to wasm_simd128.h
  • Using int64_t, uint64_t
  • More distribution of instructions (both i64 and f64)
    • comparisons (gt, ge, lt, le)
    • min/max

12 of 12

References

[0] McCutchan, John, et al. "A SIMD programming model for Dart, JavaScript, and other dynamically typed scripting languages." Proceedings of the 2014 Workshop on Programming models for SIMD/Vector processing. ACM, 2014.

Benchmark source, results, instruction at https://github.com/ngzhian/simd-benchmarks