1 of 10

Corundum status updates

Alex Forencich

2/13/2023

2 of 10

Agenda

  • Status updates

3 of 10

Status update summary

  • Bugs
    • FIFO memory inference issue (in progress)
  • Simulation updates
  • Priority flow control (todo)
  • AXI Virtual FIFO (in progress)
  • Switch version of Corundum

4 of 10

Bugs: FIFO memory inference issue

  • Seeing TX packets with incorrect IP layer checksums, only on Intel devices
  • MLE traced the issue to FIFO between TX engine and TX checksum compute block incorrectly setting “enable” bit
  • Appears to be a Quartus tool bug related to merging pipeline registers into MLABs
    • Connecting RAM output register to logic analyzer or adding “preserve” attribute results in the bug disappearing
  • Status: reported to Intel

5 of 10

Simulation updates

  • Verilator 5.006 (released Jan 22) fixes the main bug that was breaking cocotb integration
    • Verilator is much faster than Icarus Verilog and should significantly improve simulation performance
  • New issue: cannot drive non-top-level signals
    • vpi_put_value works, but value is immediately overwritten
    • Several core Corundum testbenches drive internal signals due to limitations in the signal abstractions in cocotb
    • Either need to fix this in verilator, rework cocotb to support splitting signals, or use top-level HDL testbenches to re-pack signals

6 of 10

Priority Flow Control

  • Starting to look at supporting PFC in Corundum
  • HW
    • PFC frame TX/RX, pause quanta counters
    • Per-TC queues
    • Connection to TX/RX queues and PFC frame logic
    • Internal flow control
    • TC-aware/per-TC scheduling
  • SW
    • Driver support
  • Outstanding questions
    • How to map RX traffic to priority levels (and is this necessary)?
    • How to efficiently handle multiple traffic classes in HW?
    • What needs to be done in the driver?

7 of 10

AXI virtual FIFO

  • Large packet buffer capability in DRAM
  • Store both packet data as well as sideband data
  • Intent is to support operation at 100G with all packet sizes
    • 2x DDR4-2400 channels or 2-4 HBM ports
    • Main bottleneck is memory BW, so need efficient encoding scheme for framing and sideband data
  • Status
    • Decode logic working in sim
    • Reworking FIFO channel module to support segments
    • Working on encode logic

8 of 10

Switch version of Corundum

  • Context: two different research groups interested in an FPGA-based packet switch (PFC, TSN)
    • Is anyone else interested in helping out?
  • Capabilities
    • 10G-100G operation with reasonable port count
    • Multiple traffic classes/virtual channels
    • Switching capabilities – Ethernet switching, IP routing, match/action…
    • Reasonable resource utilization
  • Target boards
    • Alveo or similar PCIe form factor – 2-4 QSFP, 2-4x 100G or 8-16x 25G
    • HTG-9200 or similar – 9 or 15 QSFP28
  • Any useful references or existing designs?

9 of 10

Switch references

  • Found several related papers from Wayne Luk/Philippos Papaphilippou
  • Feasibility investigation
    • Grouped Crosspoint Switch (GCQ) seems to work well on FPGAs
  • Hipernetch (HDL available)
    • Combined scheduling and crossbar in datapath
  • Queue-balancing switch
    • Output-queued switch with “queue balancing”
  • Outstanding questions:
    • How to reconcile packets vs. flits? (de-interleaving, drops, etc.)
    • How to implement virtual channels efficiently?

10 of 10

Switch architecture

  • Most complex parts of switch are most likely crossbar, scheduling, and route computation
  • Hierarchical switches seem to work well on FPGAs
  • Route computation is based on packet rate, not data rate
  • Share route computation logic across several ports
  • Use port groups as first level of aggregation to smaller number of higher-bandwidth crossbar ports
    • Can also mix-and-match different port groups (e.g. 4x25G + 1x100G)