OCE Virtual Chiplet eco-systems
Title: System-Level Evaluation of Multi-Vendor Chiplets Pre-RTL
Presenter: Deepak Shankar | LinkedIn
Overview of Mirabilis Design
Mirabilis Design Inc.
2
2/5/2026
Networking
RF and Analog
Why System Modeling for Chiplet Design?
Scheduling/Arbitration
proportional�share
WFQ
static
dynamic�fixed priority
EDF
TDMA
FCFS
Communication Templates
Architecture # 1
Architecture # 2
Computation Templates
DSP
ΑΙ
GPU
DRAM
CPU
FPGA
μE
DSP
TDMA
Priority
EDF
WFQ
RISC
DSP
LookUp
Cipher
ΑΙ
DSP
CPU
GPU
μE
DDR
static
Which architecture is better suited
for our application?
Workloads vs Performance- Concurrency and Contention
I/O
DSP
CPU1
CPU2
task1
task2
task3
task4
Unexpected or Expected?
System with faster Bus is slower in places
Unpredictable System Response
Spreadsheet- and Residency-based vs System Simulation
Divergence of analytical methods can lead to incorrect expectations
Using Average or Highest value
Challenges in the Design of Chiplets
CPU
CMN
DRAM
DRAM
CPU
CMN
UCIe CHI
CPU Load
1,2,4,8,16
CPU Cores
CPU
DRAM
DRAM
CPU
CPU
CMN
DRAM
DRAM
CPU
CMN
UCIe CHI
CMN
CMN
UCIe CHI
DDR Mem
System Model Requires Minimal Set of Settings
Comparing Different Configurations using UCIe Interface
All Die Adapters using PCIe 6.0
Die Adapters using PCIe 6.0 and Streaming Protocols (AXI)
Lower latency when using PCIe 6.0
Examining the Internals of UCIe IP
VisualSim UCIe IP Block Internals
UCIe System Modeling Component Verification
Note: Data verified against values in a presentation provided by Synopsys
Using the UCIe System Modeling IP in the assembly of SoC
Experiment 1 - Varying Flit_size �
Parameters | Value |
|
|
Flit_Size | 64 Bytes |
Packet_Size | 68 Bytes |
UCIe Max_Link_Speed | 32 GTPS |
Buffer Size (UCIE Tx,Rx,Retimer Rx) | 4096 |
Read Request Size | 8 Bytes |
Router_Frequency | 800MHz |
Table 1 : Base Parameters
Varying Flit_Size = Latency and Throughput across UCIe ports
Port Number | Mean Latency (μs) | Throughput (MBps) | Drop Count |
1 | 148.8 | 69.53 | 200 |
2 | 142.2 | 70.65 | 200 |
3 | 146.1 | 66.84 | 200 |
4 | - | 0.91 | 236 |
5 | 63.73 | 49.81 | 0 |
6 | 75.18 | 60.33 | 0 |
7 | 83.45 | 62.24 | 0 |
8 | 91.50 | 15.69 | 0 |
9 | 83.94 | 94.03 | 0 |
10 | 119.5 | 14.55 | 0 |
11 | 109.7 | 41.20 | 0 |
12 | 98.27 | 50.01 | 0 |
13 | 92.57 | 54.69 | 0 |
14 | 156.1 | 35.96 | 216 |
15 | 142.4 | 146.61 | 216 |
16 | 159.3 | 45.01 | 220 |
17 | 93.02 | 108.56 | 105 |
18 | - | 0.91 | 0 |
Flit Size = 64 Bytes
Port Number | Mean Latency (µs) | Throughput (MBps) | Drop Count |
1 | 117.02 | 126.44 | 134 |
2 | 104.43 | 132.09 | 118 |
3 | 91.28 | 127.51 | 96 |
4 | - | 0.91 | 236 |
5 | 117.02 | 70.88 | 0 |
6 | 28.39 | 86.67 | 0 |
7 | 53.74 | 93.77 | 0 |
8 | 43.39 | 18.79 | 0 |
9 | 44.35 | 85.51 | 0 |
10 | 49.74 | 22.51 | 0 |
11 | 35.56 | 68.16 | 0 |
12 | 32.46 | 73.88 | 0 |
13 | 23.96 | 68.81 | 0 |
14 | 134.12 | 76.73 | 175 |
15 | 146.31 | 158.95 | 180 |
16 | 136.83 | 100.12 | 183 |
17 | 74.55 | 139.88 | 57 |
18 | - | 0.91 | 0 |
Flit Size = 128 Bytes
Latency Plot for Each Chiplet
Observations on Experiment 1
Note : IO latency is measured across the PCIe, hence a difference is not observed when we vary the Flit_Size
Example Analysis
Multi-CPU Architecture
Architecture 1- Single Die with 1 Core per Cluster
SoC_with_A720AE_Demo_flow.xml
CPU
CMN
DRAM
Task Graph
Experiment 1 Results
Thermal Characteristics: Heat(J) and Temperature (°C)
With poor cooling
Time_to_Reduce_1_Degree = 1.0
With better cooling
Time_to_Reduce_1_Degree = 1.0e-8
Heat (J) is lower with better cooling
Temp (°C) is lower with better cooling
Architecture 2: Multiple Ports between Dies
SoC_with_A720AE_4Clusters_UCIe_Demo_flow.xml
CPU
CMN
DRAM
Task Graph
UCIe Links
Die 1
Die 2
Experiment 2 Results
Power has gone up per die as we have more components per die
Comparing Single vs Multi-Die Latency
OCE Virtual Chiplet eco-systems
Title: System-Level Evaluation of Multi-Vendor Chiplets Pre-RTL
Presenter: Deepak Shankar | LinkedIn