1 of 16

Comparison of FastFlow and Vitis in FPGA Stacks

for Data Centers

Rourab Paul

Paper ID : 64

Alberto Ottimo, Marco Danelutto

University of Pisa, Italy & Shiv Nadar University Chennai, India

2 of 16

Presentation Outline

  • Introduction
  • Literature Review
  • Problem formulation and proposal
  • Proposed  Method
  • Results
  • Conclusion
  • Future Work
  • References            

2

6/15/2026

3 of 16

Introduction

  • FPGA programming is more complex as compared to Central Processing Units (CPUs) and Graphics Processing Units (GPUs).
  • RTL and HLS are complex
  • It became more complex when FPGA is adopted in high-performance parallel programs in multicore platforms of data centers.
  • The unavailability of efficient high level parallel programming tools for multi core architectures makes multicore parallel programming very unpopular for the masses.

3

6/15/2026

4 of 16

Related Works

  • FastFlow is a programming library implemented in modern C++ targeting multi/many-cores and distributed systems.
  • It offers both a set of high-level ready-to-use parallel patterns and a set of mechanisms and composable components to support low-latency and high-throughput data-flow streaming networks.
  • It simplifies the development of parallel applications modelled as a structured directed graph of processing nodes.

4

6/15/2026

5 of 16

Related Works

5

6/15/2026

6 of 16

Contribution

6

6/15/2026

  • This work introduces a novel hybrid tool flow that tightly integrates FastFlow, OpenCL, and Vitis to enable efficient and scalable parallel programming for multi-FPGA stacks in data-center environments. The proposed framework automatically maps hardware kernels onto target FPGAs while guaranteeing correct and synchronized data exchange across kernel input–output interfaces.

  • The proposed approach eliminates the need for additional FPGA resources by automatically synthesizing sophisticated parallel host-side control logic directly from a lightweight csv or xml specification, significantly simplifying application development.

  • Experimental evaluation shows that hardware kernels generated using the proposed tool flow achieve performance comparable to native V itis implementations. Moreover, the framework naturally scales to large FPGA stacks and multiple concurrent hardware kernels, reducing host-side coding effort by up to ∼96% in terms of lines of code when compared with conventional Vitis-based development workflows.

7 of 16

Methods

  • This work proposes a hybrid tool flow with FastFlow, OpenCL and Vitis libraries to develop efficient and scalable parallel programs for a stack of FPGAs in data centers. This framework places hardware kernels in target FPGAs, ensuring appropriate data synchronization of their input-output ports.

7

6/15/2026

8 of 16

Methods

8

6/15/2026

9 of 16

Methods

9

6/15/2026

10 of 16

Results

10

6/15/2026

11 of 16

Results

11

6/15/2026

Example: 1

12 of 16

Realtime Task Flow

12

6/15/2026

Example: 1

13 of 16

Realtime Task Flow

13

6/15/2026

Example: 1

Example: 2

14 of 16

Conclusion & Future Works

  • Integrating FastFlow with Vitis drastically reduces the code-writing effort for parallel programming across FPGA Stacks in data centers while still achieving the same design performance as Vitis.
  • This work describes the preliminary results of our hybrid tool flow with FastFlow and Vitis libraries. The proposed script can automate the placement of hardware tasks in multiple FPGAs of data centers with appropriate data synchronization of inputs and outputs of hardware tasks.
  • This proposed framework is scalable for a stacks of FPGAs, and it automatically handles data synchronization of parallel hardware kernels running in different FPGAs.

14

6/15/2026

15 of 16

References

[1] IEEE Standards Association Corporate Advisory Group. IEEE standard for systemverilog–unified hardware design, specification, and verification language. IEEE Std 1800- 2017 (Revision of IEEE Std 1800-2012), pages 1–1315, 2018.

[2] Fabrizio Ferrandi an et al. A framework for hardware software co-design of embedded systems. Panda Home Page, 2022.

[3] Xilinx. vivado high level design features, ug1046 -ultrafast embedded design methodology guide (ug1046) (v2.3). Xilinx White Paper, 2018.

[4] AMD Xilinx. Vitis unified software platform documentation, embedded software development, ug1400 (v2022.2). Xilinx AMD White Paper, 2022.

[5] AMD. Vitis data center acceleration examples. Xilinx AMD White Paper, 2020.

[6] Marco Aldinucci, Marco Danelutto, Peter Kilpatrick, and Massimo Torquati. Fastflow: High-Level and Efficient Streaming on Multicore, chapter 13, pages 261–280. John Wiley & Sons, Ltd, 2017.

[7] Nicolò Tonci, Massimo Torquati, Gabriele Mencagli, and Marco Danelutto. Distributed-memory fastflow building blocks. International Journal of Parallel Programming, 51(1):1–21, Feb 2023.

[8] Marco Danelutto, Gabriele Mencagli, Alberto Ottimo, Francesco Iannone, and Paolo Palazzari. Fastflow targeting

15

6/15/2026

16 of 16

THANK YOU