1 of 20

Breaking the Memory Capacity-Bandwidth Tradeoff for Agentic Inference

Allan Cantle

CEO, Nallasway

Draft V1.1 - 9/1/2026

Contributors : Amphenol, Crealien, Credo, Empower(Analog Devices), FIT, LOTES

SERVER: HPC

2 of 20

Why Inference Keeps Eating Memory

More tokens drive bandwidth · bigger models + longer context drive capacity

1

Bigger models — MoE

10× total params, same active set

CAPACITY

2

Long context + agents

KV cache grows with every token

CAPACITY

BANDWIDTH

3

Reasoning tokens

10–100× tokens per answer

BANDWIDTH

Every token re-reads it all —

capacity drags bandwidth with it.

Memory per serving instance

KV cache now rivals the weights

0.4 TB

1.2 TB

5 TB

Illustrative, order-of-magnitude

3 of 20

Credo’s Lightweight Serial Interconnect, LSI, OCP Contribution

  • OmniConnect LSI
    • Unified XPU IO Chiplet Interface
      • Late IO Binding, e.g. X% Memory, Y% Scale-up, Z% Scale-Out
    • AXI over VSR 112Gbps Serdes @ 1.15pJ/bit
    • Low Latency, ~35ns one way
  • Compliant with OCP FCSA Section 5.12
    • Non-Coherent Memory Traffic Interface
  • 600 Serdes East/West IO ~8 TBytes/s
    • Per Reticule Sized Die

4 of 20

Bandwidth per Reticle-Sized Die

Jalapeño went back to one reticle die with six HBM4 stacks; Rubin Ultra needs early HBM4E to stay ahead

Package figures from SemiAnalysis. Rubin Ultra 2027 is the revised two-die / 8-stack part (four-die version cancelled, June 2026), plotted on Samsung HBM4E — 16 Gbps, 4.0 TB/s per stack, GTC 2026 — assuming early HBM4E adoption; a mainstream HBM4 8-Hi SKU is also reported at ~21 TB/s per package. Jalapeño from Hot Chips 2026 — 15.4 TB/s, 216 GiB HBM4, 700 W.

5 of 20

Bandwidth per Reticle-Sized Die

Weaver-2 reaches 11.5 TB/s per die — 72% of Rubin Ultra, with no HBM and no advanced packaging

Package figures from SemiAnalysis. Rubin Ultra 2027 is the revised two-die / 8-stack part (four-die version cancelled, June 2026), plotted on Samsung HBM4E — 16 Gbps, 4.0 TB/s per stack, GTC 2026 — assuming early HBM4E adoption; a mainstream HBM4 8-Hi SKU is also reported at ~21 TB/s per package. Jalapeño from Hot Chips 2026 — 15.4 TB/s, 216 GiB HBM4, 700 W. Weaver figures are Nallasway targets.

6 of 20

Capacity per Reticle-Sized Die

Per-die capacity growth broke at Blackwell — 36%/yr through 2023, 8%/yr since, and nobody is back on trend

Package figures from SemiAnalysis. Rubin Ultra 2027 on 12-high HBM4E — 48 GB per stack, four per reticle die — same early-HBM4E basis as the bandwidth slide; a mainstream HBM4 8-Hi SKU is also reported at 192 GB per package. Jalapeño: 216 GiB HBM4, six stacks, one die (Hot Chips 2026).

7 of 20

Capacity per Reticle-Sized Die

Weaver puts per-die capacity back above trend — 3.2× the HBM part at minimum, 25× at maximum

Package figures from SemiAnalysis. Rubin Ultra 2027 on 12-high HBM4E — 48 GB per stack, four per reticle die — same early-HBM4E basis as the bandwidth slide; a mainstream HBM4 8-Hi SKU is also reported at 192 GB per package. Jalapeño: 216 GiB HBM4, six stacks, one die (Hot Chips 2026). Weaver range is the build-time module choice — 50 × 16–64 GB, 76 × 8–32 GB.

8 of 20

Three ways to attach memory today

CPU + 16 MRDIMMs

DDR5 MRDIMM-8800 · 256 GB per DIMM

Bounding Area ≈ 330 cm²

4 TB

1.1 TB/s

CAPACITY

BANDWIDTH

Capacity, not bandwidth

CPU

Vera + 8 SOCAMMs

SOCAMM2 LPDDR5X-9600 · up to 48 GB each

Bounding Area ≈ 178 cm²

1.5 TB

1.2 TB/s

CAPACITY

BANDWIDTH

Compact, but capped

Vera

Reticule XPU + 50 OC-DDIMMs

LPDDR5X-9600 OC-DDIMM - 64GB per DIMM (OmniConnect Differential DIMMs)

XPU

Bounding Area = 138 cm²

3.2 TB

7.5 TB/s

CAPACITY

BANDWIDTH

One Memory Tier

​

Delivers Compact Capacity at Bandwidth

50 mm

Jalapeño + 6 HBM4

HBM4 · 36GB each

Bounding Area = 65 cm²

216 GB

15.4 TB/s

CAPACITY

BANDWIDTH

Bandwidth, not capacity

9 of 20

OmniConnect Weaver 1 Memory Buffer & Module

  • Weaver 1 Memory Buffer
    • Supports 4 LPDDDR5X x32 Channels
      • 9.6Gbps Data Rate
    • 12x 112Gb Transceivers (3 per Channel)
  • Proposed Weaver 1 OC-DDIMM Module
    • 2x LPDDR5X x16 4 Ch 563b packages
    • 16 GByte to 64 GByte Capacity
      • 4 to 8, 16Gb or 32Gb Die Stacks per package
  • OC-DDIMM Connector
    • Based on M.2 Connector
    • 96 pads at 0.5mm pitch on a 0.8mm thick Substrate

Weaver 1 Memory Buffer

29mm

24mm

4mm

10 of 20

OC-DDIMM Connector, M.2 Derivative, Design & PCIe-G7 SI

5.5mm

0.8mm

3.5mm

29mm

0.5mm

Current In : 1A per Pin @3.3V

Total 6 Pins supporting 20W

OC-DDIMM 96 Pin Connector

11 of 20

OCP HPCM Module — Block Diagram

50 buffered-LPDDR memory modules over 600 × 112G OmniConnect lanes · 128 × fabric lanes

Weaver Memory Module

16 GB – 64 GB

Weaver Memory Module

16 GB – 64 GB

12 × 112G

transceivers

12 × 112G

transceivers

XPU

Full Reticule

Compute Die

128 Transceivers (112G OmniConnect or 224G LR)

Scale Up / Scale Out Fabric

Via LOTES PrezLink CPC

25 × Socketed

Memory Modules

25 × Socketed

Memory Modules

Front module shown in full; 24 identical modules staggered behind it on each side. 12 × 112G per module → 300 lanes per side, 600 lanes total. Capacity 0.8 TB – 3.2 TB per XPU module.

12 of 20

OCP HPCM 2026 Module - 50 Memory Modules + CPC IO

  • Features
    • Standalone HPC XPU Module
      • No Motherboard required
    • 128 SerDes Co-Packaged Copper IO
      • LOTES PrezLink
    • Organic SubStrate with Monolithic Die
    • Hybrid Single / 2 Phase Water Cooled
    • Socketed CoPackaged Memory DIMMs
      • 50 Weaver Memory Modules
    • 48V Power Input, 1,650W TDP

115mm

150mm

40mm

LOTES

PrezLink

X8

13 of 20

OCP HPCM 2026 Module - 50 Memory Modules + CPC IO

  • Vertical Backside Power Delivery
    • LuxShare
      • 2x Unregulated 48V to 3.3V Modules
      • 250A Output per Module
    • EMPOWER (now part of Analog Devices)
      • 24x EP7502MC 65A POL Modules
  • EP7144 used on OC-DDIMM Modules

115mm

150mm

40mm

14 of 20

OCP HPCM 2026 Module - 50 Memory Modules + CPC IO

  • Socketed CoPackaged Memory DIMMs
    • No Module Assembly Yield Loss
    • Serviceable
      • Replace Failing Memories
  • All Interconnect On Substrate
    • 1.15pJ/bit for all XPU IO
    • Including between 18 XPUs in a 2OU Node

15 of 20

OC-DDIMM Routability Study

  • 600 Diff Pairs routed from 33mm Reticule Die Edge to OC-DDIMMs
  • 26um Trace Width
    • Relaxed for Large Substrate Yield
    • 15um Trace for Die Breakout
  • 7.5dB Worst Case substrate routing Loss
    • XPU Die Bump to Connector Pad
  • Total Channel OIF VSR Budget ~25dB?
    • Including Die Bump to Package Ball
    • Plenty of Margin

16 of 20

ColdPlate Design Architecture - with 2 Phase Water Cooling

  • Hybrid Single/Two-Phase Architecture
    • Divides the cold plate into a Single-Phase Section (Zones 1–3) for DIMM preheating and a Two-Phase Section (Zones 4–5) for XPU latent heat dissipation.
  • Cumulative Upstream Hydraulic Restrictor Strategy
    • The total cumulative single-phase pressure drop across Zones 1–3 acts as an integrated upstream throttling barrier. Suppress boiling back-pressure and preventing flow reversal.
  • Sub-Atmospheric Thermal Clamping
    • Under an exit vacuum pressure of -90kPa, pure water boils at ~48C-53C, maintaining XPU junction temperatures in an optimal range.

Zone1

Inlet

Zone2

DIMM

Zone 4

XPU

Zone5

Outlet

Single Phase

Two Phase

CreAlien Co Ltd

Zone 3

XPU

17 of 20

Cold Plate Design

  • CFD Flow Boundary & Power Capacity Target
    • Current CFD models a peak flow condition of 2 LPM total (1 LPM per side)
    • Based on empirical data (0.4 LPM / kW), this boundary evaluates hydraulic pressure limits for high-power cooling up to ~5 kW XPU load.
  • Symmetric Flow Balancing & Uniform Distribution
    • Verifies balanced dual-inlet split flow across left and right DIMM micro-fin banks, ensuring uniform velocity and pressure drop before merging into the central core.
  • Next Actions: Multi-Regime & Thermal Coupling Analysis
    • Low-Flow / Low-Power Dynamics: Evaluate system hydraulic behavior at reduced flow rates to analyze alternative flow regimes and stability limits.
    • Thermal Source Coupling: Integrate heat source inputs—prioritizing single-phase DIMM heat flux—to evaluate real-time fluid preheating and saturation point shifts.

CreAlien Co Ltd

18 of 20

16 to 18 HPCM2026 Modules in an ORV3 2OU Chassis

  • Power & Water Cooling provided from below Modules
  • Modules Blind Mate into Power and Water
  • 30KW Power Budget, 1.65KW per Module
    • 800W XPU
    • 850W Memory Modules

19 of 20

Call to Action - Help us bring this HPCM2026 Concept to POC

  • We need help to bring this concept to reality
    • Especially adopters & Interested parties in Credo’s OmniConnect Technology
    • Help create a compelling PoC that can grow into a competitive viable ecosystem
  • Join our OCP HPC Subproject Workgroup
    • Mailing list: https://ocp-all.groups.io/g/OCP-HPC
    • Wiki : https://www.opencompute.org/wiki/HPC
    • Meeting Calendar : https://www.opencompute.org/projects/project-and-ic-meetings-calendar
    • Every other Tuesday, 8am Pacific

20 of 20

Thank You!