1 of 8

InfiniBand bandwidth and latency comparison in bare-metal and virtualised environments

Jacob Anders for OpenStack Scientific SIG

23 January 2019

2 of 8

The goals

  • Establish performance baseline across core resources,
  • Emphasise the value of bare-metal Cloud,
  • Gain better understanding of why people complain about the performance of virtualised MPI applications,
  • Help decide which applications need bare-metal and which can be considered for consolidation in performance optimised VMs

3 of 8

Test environment

  • Test environment is described in the table below:

  • We stuck as close to “reasonable defaults” as possible - enabled performance mode in BIOS and CPU passthrough in nova, but didn’t try CPU pinning, recompiling software to maximise performance - we wanted to remain as generic as possible

4 of 8

Results

  • ib_write_bw uses 5000 iterations and 65536 byte messages,
  • ib_write_lat uses 2 byte messages and 1000 iterations

5 of 8

Relative comparison of bare-metal and VMs

  • Here we look at relative performance as well

6 of 8

Relative comparison of bare-metal and VMs

  • Now we compare standard deviation, not absolute numbers

7 of 8

Conclusions

  • Bandwidth in SRIOV/IB VMs can match bare-metal,
  • Latency for small message sizes, on the other hand, can not,
  • Interestingly, variability of the test is an order of magnitude higher in VMs than in bare-metals
  • This WILL likely upset the MPI apps
  • Future work: how would CPU pinning affect:
    • Latency
    • Bandwidth
    • Linpack! (it would likely cost us cores!)
  • Overheads on the Linpack seem higher than what we’ve seen in the past - impact of Spectre/Meltdown fixes?
  • It will be very interesting to re-test on newer hardware and compare (shipping soon)

8 of 8

Thank you

  • Questions?