---
title: 'LaMoSys3.5D: Heterogeneous 3.5D-IC for LLM Inference'
url: https://www.emergentmind.com/topics/lamosys3-5d
type: topic
---

# LaMoSys3.5D: Heterogeneous 3.5D-IC for LLM Inference

LaMoSys3.5D is a scalable 3.5D integrated circuit (IC) architecture designed to optimize large language model (LLM) inference serving through a heterogeneous composition of 3D-DRAM chiplets on a 2.5D silicon interposer. This platform employs a hardware/software co-design paradigm to balance the compute-intensive prefill and the bandwidth-intensive decode phases, providing improved throughput-per-watt and significantly lower latency relative to state-of-the-art GPU and prior 3D-DRAM-based inference systems [2512.08731].

## 1. Heterogeneous 3.5D-IC Architecture

LaMoSys3.5D organizes heterogeneous chiplets into a grid, leveraging two principal types:

- **Prefill-optimized Chiplets (PC):** These are compute-rich, containing higher counts of processing elements (PEs) and fewer DRAM layers ($N_{\rm dram}=4$).
- **Decode-optimized Chiplets (DC):** Bandwidth- and capacity-rich, these chiplets utilize more DRAM layers ($N_{\rm dram}=8$), wider through-silicon-via (TSV) interfaces, larger on-die DRAM capacity, and fewer PEs.

A representative system instantiates a $2\times2$ or $3\times3$ chiplet grid, partitioning PCs and DCs to spatially align with the distinct traffic profiles of model weights, activations, and the key/value (KV) cache. The chiplets are interconnected by four AIB-2.0 PHY ports, each supporting up to 200 GB/s, yielding 800 GB/s aggregate chiplet interconnect bandwidth. Internally, each chiplet features a mesh network-on-chip (NoC), with flit widths attuned to DRAM channel widths.

Each chiplet stacks DRAM atop a 7 nm logic die via hybrid bonding, with each DRAM layer comprising $N_{\rm bank}$ banks and TSV interfaces $N_{\rm io}$ bits wide (typ. $N_{\rm bank}\in\{16,32,64\}$, $N_{\rm io}\in\{128,256,512\}$). DRAM command scheduling uses closed-page, FCFS policy, and layer-interleaved bank distribution to mask activation and transfer timing ($t_{\rm RCD}$, $t_{\rm CAS}$, $t_{\rm RP}$; JESD79-4).

## 2. Dataflow and Parallelization Co-Design

### 2.1 Intra-PE Dataflow: Direct-DRAM-Delivery (D³)

LaMoSys3.5D introduces the "D³" intra-PE dataflow where GEMM computations $C \leftarrow A \times B + C$ are tiled by $(T_M,T_N,T_K)$ and assigned a reuse policy $RU \in \{\mathrm{IRU}, \mathrm{WRU}, \mathrm{ORU}, \mathrm{ARU}\}$ subject to on-PE SRAM buffer size $S_{\rm buf}$. Feasibility conditions are imposed (e.g., input-reuse: $T_M T_K \leq S_{\rm buf}$). The D³ mechanism enables selective tiles (such as $B$ weights) to be directly streamed from 3D-DRAM, with critical tiles staged in SRAM. Analytical cost modeling for a given tile configuration estimates execution cycles as follows:

\[
L_{\rm tile} = L_{\rm compute}(T_M,T_N,T_K) + L_{\rm mem}(DATA_{\rm from\,DRAM}) + L_{\rm comm}
\]

An exhaustive search across the design space yields the (nearly) optimal tile and reuse policy.

### 2.2 PE Mapping and Scheduling

Parallelization employs tensor-parallel (TP), pipeline-parallel (PP), and optional data-parallel (DP) schemes. TP groups of PEs are determined per pipeline stage $i$ via a mixed-integer linear program (MILP) to minimize both all-reduce diameter and inter-group distances:

\[
\sum_i \Bigl(\text{diameter}(G_i) + w_{\rm inter} \sum_{\langle i,j \rangle} \| \mathrm{center}_i - \mathrm{center}_j \| \Bigr)
\]

Simulated annealing is used for PP stage placement to further minimize stage latency and KV handoff costs. A dynamic ORCA-style scheduler overlaps communication and adjacent compute to maximize throughput for dominant transformer computations (QKV projections, output projections, feedforward layers).

## 3. Thermal-Aware Modeling and Hierarchical Design Space Exploration

### 3.1 Thermal Modeling

A compact electrical circuit model represents thermal resistances ($R_{\rm th}$) in the 3.5D stack, covering DRAM layers, logic, and cooling path. The composite temperature is given by:

\[
T_{\max} = T_{\rm amb} + \bigl( P_{\rm leak}(T) + P_{\rm dyn} \bigr) R_{\rm th}
\]

where leakage power $P_{\rm leak}(T)$ increases roughly exponentially, and DRAM refresh interval $t_{\rm RFI}$ halves every $10^{\circ}$C increase above $85^{\circ}$C, reducing accessible bandwidth. Transient temperature evolution is simulated using ATSim3D, and DRAM refresh penalties are dynamically accounted for in memory performance.

### 3.2 Hierarchical Design Space Exploration (DSE)

At the chiplet level, the design space is spanned over $N_{\rm dram}$, $N_{\rm bank}$, $N_{\rm io}$, number of cores, base-SA partition sizes, and SRAM capacities. Bayesian optimization synthesizes Pareto-efficient solutions for compute (TFLOPS), bandwidth (TB/s), and memory capacity (GB). At the system level, permutations of PC/DC chiplet ratios are evaluated to maximize throughput-per-watt, subject to strict service-level objectives (SLOs), power, thermal, area, and capacity constraints such as:

\[
\max_{\text{design}} \quad \frac{\text{Throughput}}{\text{Watt}}
\quad \text{s.t.} \quad
\mathrm{TTFT} \leq \mathrm{TTFT}_{\max},~
\mathrm{TBT} \leq \mathrm{TBT}_{\max},~
T_{\max} \leq T_{\rm lim},~
P_{\rm peak} \leq P_{\rm rack},~
\mathrm{Capacity} \geq \mathrm{KV\_size},~
\text{area} \leq \text{interposer\_area},~
\text{pins} \leq \text{pin\_budget}
\]

Simulation-based event-driven loops converge on designs satisfying these multidimensional criteria.

## 4. Quantitative Performance Analysis

Comprehensive benchmarking of LaMoSys3.5D highlights substantial improvements over baseline and prior platforms:

| Metric                      | LaMoSys3.5D                    | DGX-A100 / 3D Baselines       |
|-----------------------------|---------------------------------|-------------------------------|
| Throughput-per-Watt         | 0.75 tokens/s/W                 | 0.46 tokens/s/W (+62% LaMoSys)|
| Prefill TTFT (Time-to-First)| 2.13× DGX-A100                  | -                             |
| Decode TBT (Time/Token)     | up to 17× faster                | -                             |
| End-to-end latency (geo-mean)| 4.87× lower than single-3D      | TETRIS: 2.99×, 3D-LC: 3.02×, 3D-TokSIM: 8.58×  |
| Decode EDP (batch=16, len=2048) | D³ baseline                   | Token-stationary: +13%, ARU: +9%, TETRIS: +1%  |

For prefill (compute-bound), performance deltas are minor ($\leq 3\%$), while decode (memory-bound) demonstrates D³'s efficiency. Under benchmarked traces, PC maximum temperatures reach $\approx100^{\circ}$C, DC reach $\approx95^{\circ}$C; a $40^{\circ}$C temperature rise reduces DRAM bandwidth by $\sim 10\%$ and increases logic leakage by $\sim 20\%$.

## 5. Design Principles and Practical Guidelines

Key principles distilled from LaMoSys3.5D's evaluation include:

1. **PD-disaggregation on 3.5D-ICs:** Strategic co-location of compute-lean PCs and memory-abundant DCs enables prefill vs. decode workload alignment.
2. **Short-wide base-SAs + D³ flow:** Systolic arrays are subdivided for high GEMV utility at low batch, permitting direct weight streaming from DRAM.
3. **Exhaustive intra-PE mapping + mesh-aware PE placement:** Full (TM, TN, TK, RU) enumeration and mesh-centric MILP/SA grouping minimize communication overheads.
4. **Coupled thermal-performance modeling:** Integrating transient thermal simulation with refresh-aware memory models precludes bandwidth erosion at elevated stack temperatures; core counts per chiplet must be limited ($\leq 16$ cores/PE) to respect thermal budgets.
5. **Hierarchical and constraint-driven DSE:** Chiplet vs. system design objectives are cleanly partitioned and jointly solved under QoS, area, power, and capacity constraints.

## 6. Significance and Impact on Inference Serving

LaMoSys3.5D demonstrates, for the first time, a platform in which a 3.5D-IC with heterogeneous 3D-DRAM chiplets, DRAM-native dataflow, mesh-aware parallel mapping, and early thermal modeling collectively realizes $>60\%$ increase in token/W and $4{-}9\times$ latency reduction relative to the best-known GPU- and 3D-DRAM-based inference architectures. This represents a significant evolution in balancing architectural specialization with end-to-end LLM serving efficiency [2512.08731].

Source: https://www.emergentmind.com/topics/lamosys3-5d