1 of 25

OCE Virtual Chiplet eco-systems

Title: System-Level Evaluation of Multi-Vendor Chiplets Pre-RTL

Presenter: Deepak Shankar | LinkedIn

2 of 25

Overview of Mirabilis Design

  • EDA Software Company based in Silicon Valley
    • Rrisk assessment, performance optimization and validation of specification
  • Highly experience Management and Engineering team
    • Over 150 man-years in semiconductors, automotive and Automotive
  • VisualSim Architect –Design the Right product
    • Simulation with System Modeling Components for Architecture Exploration
  • Shift-Left to Shift-Right- A Total Solution
    • System Optimization, Collaboration, Validation and Early Reference Designs

​

​

​

​

Mirabilis Design Inc.

2

2/5/2026

Networking

RF and Analog

3 of 25

Why System Modeling for Chiplet Design?

Scheduling/Arbitration

proportional�share

WFQ

static

dynamic�fixed priority

EDF

TDMA

FCFS

Communication Templates

Architecture # 1

Architecture # 2

Computation Templates

DSP

ΑΙ

GPU

DRAM

CPU

FPGA

μE

DSP

TDMA

Priority

EDF

WFQ

RISC

DSP

LookUp

Cipher

ΑΙ

DSP

CPU

GPU

μE

DDR

static

Which architecture is better suited

for our application?

4 of 25

Workloads vs Performance- Concurrency and Contention

I/O

DSP

CPU1

CPU2

task1

task2

task3

task4

  • Contention�- limited resources�- scheduling/arbitration

​

  • Interference of multiple applications�- limited resources�- scheduling/arbitration�- anomalies
  • Complex behavior�- input stream�- data dependent behavior

​

5 of 25

Unexpected or Expected?

System with faster Bus is slower in places

Unpredictable System Response

6 of 25

Spreadsheet- and Residency-based vs System Simulation

Divergence of analytical methods can lead to incorrect expectations

Using Average or Highest value

7 of 25

Challenges in the Design of Chiplets

  • Latency, bandwidth and power consumption changes when transition from single-die to multi-die
  • Impact of peripheral, memories and system cache position
  • Scalable to support current and future applications?
  • Support sufficient ports for application-specific topology?
  • Distribution of accelerators and scheduling of tasks

CPU

CMN

DRAM

DRAM

CPU

CMN

UCIe CHI

CPU Load

1,2,4,8,16

CPU Cores

CPU

DRAM

DRAM

CPU

CPU

CMN

DRAM

DRAM

CPU

CMN

UCIe CHI

CMN

CMN

UCIe CHI

DDR Mem

8 of 25

System Model Requires Minimal Set of Settings

9 of 25

Comparing Different Configurations using UCIe Interface

All Die Adapters using PCIe 6.0

Die Adapters using PCIe 6.0 and Streaming Protocols (AXI)

Lower latency when using PCIe 6.0

10 of 25

Examining the Internals of UCIe IP

11 of 25

VisualSim UCIe IP Block Internals

12 of 25

UCIe System Modeling Component Verification

Note: Data verified against values in a presentation provided by Synopsys

13 of 25

Using the UCIe System Modeling IP in the assembly of SoC

  • Core is an UCIe interface Module
  • 4 processor Chiplets and 1 I/O Chiplet
    • Each Processor Chiplet uses 4 UCIe Links.
    • The I/O Chiplet uses 1 UCIe link.
  • Workload
    • Each Processor Chiplet has 4 traffic generators
    • Send packets to the NOC UCIe block
  • Transactions data flow:
    • CPU -- > DSP
    • DSP -- > GPU
    • GPU 🡪 AI
    • AI 🡪 I/O and CPU
    • I/O 🡪 GPU

​

14 of 25

Experiment 1 - Varying Flit_size �

Parameters

Value

 

 

Flit_Size

64 Bytes

Packet_Size

68 Bytes

UCIe Max_Link_Speed

32 GTPS

Buffer Size (UCIE Tx,Rx,Retimer Rx)

4096

Read Request Size

8 Bytes

Router_Frequency

800MHz

Table 1 : Base Parameters

15 of 25

Varying Flit_Size = Latency and Throughput across UCIe ports

Port Number

Mean Latency (μs)

Throughput (MBps)

Drop Count

1

148.8

69.53

200

2

142.2

70.65

200

3

146.1

66.84

200

4

-

0.91

236

5

63.73

49.81

0

6

75.18

60.33

0

7

83.45

62.24

0

8

91.50

15.69

0

9

83.94

94.03

0

10

119.5

14.55

0

11

109.7

41.20

0

12

98.27

50.01

0

13

92.57

54.69

0

14

156.1

35.96

216

15

142.4

146.61

216

16

159.3

45.01

220

17

93.02

108.56

105

18

-

0.91

0

Flit Size = 64 Bytes

Port Number

Mean Latency (µs)

Throughput (MBps)

Drop Count

1

117.02

126.44

134

2

104.43

132.09

118

3

91.28

127.51

96

4

-

0.91

236

5

117.02

70.88

0

6

28.39

86.67

0

7

53.74

93.77

0

8

43.39

18.79

0

9

44.35

85.51

0

10

49.74

22.51

0

11

35.56

68.16

0

12

32.46

73.88

0

13

23.96

68.81

0

14

134.12

76.73

175

15

146.31

158.95

180

16

136.83

100.12

183

17

74.55

139.88

57

18

-

0.91

0

Flit Size = 128 Bytes

16 of 25

Latency Plot for Each Chiplet

17 of 25

Observations on Experiment 1

  • Latency Reduces Significantly – both Per Chiplet and across the UCIe ports when Flit_Size is doubled
  • Throughput across the UCIe ports Increases Significantly - when Flit_Size is doubled
  • Number of packets dropped also reduces when we increase the Flit_Size

​

Note : IO latency is measured across the PCIe, hence a difference is not observed when we vary the Flit_Size

18 of 25

Example Analysis

Multi-CPU Architecture

19 of 25

Architecture 1- Single Die with 1 Core per Cluster

SoC_with_A720AE_Demo_flow.xml

CPU

CMN

DRAM

Task Graph

20 of 25

Experiment 1 Results

  • Beh_Flow_Latency – Shows the time it took to complete a task by CPU clusters (Chiplet) 1 and 2
  • Latency plot – Shows the time taken by memory access requests sent out to the DRAM to complete
  • Power_Plot – shows the power consumption across the die
  • Text statistics – Shows detailed statistics across all cache levels (all chiplets), buses and DRAM

21 of 25

Thermal Characteristics: Heat(J) and Temperature (°C)

With poor cooling

Time_to_Reduce_1_Degree = 1.0

With better cooling

Time_to_Reduce_1_Degree = 1.0e-8

Heat (J) is lower with better cooling

Temp (°C) is lower with better cooling

22 of 25

Architecture 2: Multiple Ports between Dies

SoC_with_A720AE_4Clusters_UCIe_Demo_flow.xml

CPU

CMN

DRAM

Task Graph

UCIe Links

Die 1

Die 2

23 of 25

Experiment 2 Results

  • Memory_Distribution_Percentage to increase DRAM accesses to other DRAM’s
  • Increase the number of Cores per cluster from 1 to 16
  • Experiment Routing Table to create the best route with minimal bottleneck

​

​

Power has gone up per die as we have more components per die

24 of 25

Comparing Single vs Multi-Die Latency

  • Comparing Memory address distribution = 0% vs 40.0%
  • Increase the address distribution percentage
    • Task latency increases as data from DRAM in another die

25 of 25

OCE Virtual Chiplet eco-systems

Title: System-Level Evaluation of Multi-Vendor Chiplets Pre-RTL

Presenter: Deepak Shankar | LinkedIn