Breaking the Memory Capacity-Bandwidth Tradeoff for Agentic Inference
Allan Cantle
CEO, Nallasway
Draft V1.1 - 9/1/2026
Contributors : Amphenol, Crealien, Credo, Empower(Analog Devices), FIT, LOTES
SERVER: HPC
Why Inference Keeps Eating Memory
More tokens drive bandwidth · bigger models + longer context drive capacity
1
Bigger models — MoE
10× total params, same active set
CAPACITY
2
Long context + agents
KV cache grows with every token
CAPACITY
BANDWIDTH
3
Reasoning tokens
10–100× tokens per answer
BANDWIDTH
Every token re-reads it all —
capacity drags bandwidth with it.
Memory per serving instance
KV cache now rivals the weights
0.4 TB
1.2 TB
5 TB
Illustrative, order-of-magnitude
Credo’s Lightweight Serial Interconnect, LSI, OCP Contribution
Bandwidth per Reticle-Sized Die
Jalapeño went back to one reticle die with six HBM4 stacks; Rubin Ultra needs early HBM4E to stay ahead
Package figures from SemiAnalysis. Rubin Ultra 2027 is the revised two-die / 8-stack part (four-die version cancelled, June 2026), plotted on Samsung HBM4E — 16 Gbps, 4.0 TB/s per stack, GTC 2026 — assuming early HBM4E adoption; a mainstream HBM4 8-Hi SKU is also reported at ~21 TB/s per package. Jalapeño from Hot Chips 2026 — 15.4 TB/s, 216 GiB HBM4, 700 W.
Bandwidth per Reticle-Sized Die
Weaver-2 reaches 11.5 TB/s per die — 72% of Rubin Ultra, with no HBM and no advanced packaging
Package figures from SemiAnalysis. Rubin Ultra 2027 is the revised two-die / 8-stack part (four-die version cancelled, June 2026), plotted on Samsung HBM4E — 16 Gbps, 4.0 TB/s per stack, GTC 2026 — assuming early HBM4E adoption; a mainstream HBM4 8-Hi SKU is also reported at ~21 TB/s per package. Jalapeño from Hot Chips 2026 — 15.4 TB/s, 216 GiB HBM4, 700 W. Weaver figures are Nallasway targets.
Capacity per Reticle-Sized Die
Per-die capacity growth broke at Blackwell — 36%/yr through 2023, 8%/yr since, and nobody is back on trend
Package figures from SemiAnalysis. Rubin Ultra 2027 on 12-high HBM4E — 48 GB per stack, four per reticle die — same early-HBM4E basis as the bandwidth slide; a mainstream HBM4 8-Hi SKU is also reported at 192 GB per package. Jalapeño: 216 GiB HBM4, six stacks, one die (Hot Chips 2026).
Capacity per Reticle-Sized Die
Weaver puts per-die capacity back above trend — 3.2× the HBM part at minimum, 25× at maximum
Package figures from SemiAnalysis. Rubin Ultra 2027 on 12-high HBM4E — 48 GB per stack, four per reticle die — same early-HBM4E basis as the bandwidth slide; a mainstream HBM4 8-Hi SKU is also reported at 192 GB per package. Jalapeño: 216 GiB HBM4, six stacks, one die (Hot Chips 2026). Weaver range is the build-time module choice — 50 × 16–64 GB, 76 × 8–32 GB.
Three ways to attach memory today
CPU + 16 MRDIMMs
DDR5 MRDIMM-8800 · 256 GB per DIMM
Bounding Area ≈ 330 cm²
4 TB
1.1 TB/s
CAPACITY
BANDWIDTH
Capacity, not bandwidth
CPU
Vera + 8 SOCAMMs
SOCAMM2 LPDDR5X-9600 · up to 48 GB each
Bounding Area ≈ 178 cm²
1.5 TB
1.2 TB/s
CAPACITY
BANDWIDTH
Compact, but capped
Vera
Reticule XPU + 50 OC-DDIMMs
LPDDR5X-9600 OC-DDIMM - 64GB per DIMM (OmniConnect Differential DIMMs)
XPU
Bounding Area = 138 cm²
3.2 TB
7.5 TB/s
CAPACITY
BANDWIDTH
One Memory Tier
Delivers Compact Capacity at Bandwidth
50 mm
Jalapeño + 6 HBM4
HBM4 · 36GB each
Bounding Area = 65 cm²
216 GB
15.4 TB/s
CAPACITY
BANDWIDTH
Bandwidth, not capacity
OmniConnect Weaver 1 Memory Buffer & Module
Weaver 1 Memory Buffer
29mm
24mm
4mm
OC-DDIMM Connector, M.2 Derivative, Design & PCIe-G7 SI
5.5mm
0.8mm
3.5mm
29mm
0.5mm
Current In : 1A per Pin @3.3V
Total 6 Pins supporting 20W
OC-DDIMM 96 Pin Connector
OCP HPCM Module — Block Diagram
50 buffered-LPDDR memory modules over 600 × 112G OmniConnect lanes · 128 × fabric lanes
Weaver Memory Module
16 GB – 64 GB
Weaver Memory Module
16 GB – 64 GB
12 × 112G
transceivers
12 × 112G
transceivers
XPU
Full Reticule
Compute Die
128 Transceivers (112G OmniConnect or 224G LR)
Scale Up / Scale Out Fabric
Via LOTES PrezLink CPC
25 × Socketed
Memory Modules
25 × Socketed
Memory Modules
Front module shown in full; 24 identical modules staggered behind it on each side. 12 × 112G per module → 300 lanes per side, 600 lanes total. Capacity 0.8 TB – 3.2 TB per XPU module.
OCP HPCM 2026 Module - 50 Memory Modules + CPC IO
115mm
150mm
40mm
LOTES
PrezLink
X8
OCP HPCM 2026 Module - 50 Memory Modules + CPC IO
115mm
150mm
40mm
OCP HPCM 2026 Module - 50 Memory Modules + CPC IO
OC-DDIMM Routability Study
ColdPlate Design Architecture - with 2 Phase Water Cooling
Zone1
Inlet
Zone2
DIMM
Zone 4
XPU
Zone5
Outlet
Single Phase
Two Phase
CreAlien Co Ltd
Zone 3
XPU
Cold Plate Design
CreAlien Co Ltd
16 to 18 HPCM2026 Modules in an ORV3 2OU Chassis
Call to Action - Help us bring this HPCM2026 Concept to POC
Thank You!