---
title: 'Punica: Pomegranate Propagation & LLM Serving'
url: https://www.emergentmind.com/topics/punica
type: topic
---

# Punica: Pomegranate Propagation & LLM Serving

Punica appears in the cited literature in two distinct technical senses. In plant science and microscopy, it denotes *Punica granatum* (pomegranate), examined both as hardwood cuttings in propagation experiments and as pollen cells in high-resolution 3D morphology reconstruction. In systems research, “Punica” denotes a multi-tenant LoRA serving system for large language models that exploits a shared pretrained backbone, a new CUDA kernel, and a consolidation-oriented scheduler for efficient GPU-cluster operation [2311.04953][2407.00408][2506.07145][2310.18547].

## 1. Scope and research usages

The current literature represented here uses the name across biological and computational domains rather than in a single uniform sense. In the biological papers, the emphasis is on vegetative propagation and cellular morphology of *Punica granatum*. In the systems paper, the same name identifies an LLM-serving framework for multi-tenant LoRA workloads.

| Usage of “Punica” | Object studied | Representative result |
|---|---|---|
| *Punica granatum* in propagation | Hardwood cuttings | Best interaction treatment: **49.99%** rooting with **straight basal cut + 1 cm or 2 cm basal internode stub** |
| *Punica granatum* in microscopy | Pollen cell | Reconstructed as **roughly spherical** with **three prominent protrusions**; volume **6021.3 μm³** |
| Punica as a systems framework | Multi-tenant LoRA serving | **12x higher throughput** with only **2 ms latency per token** |

This multiplicity of usage is central to interpreting the term in recent technical writing: in one branch it functions as a biological specimen designation, and in another as a proper name for an inference-serving system.

## 2. Hardwood cutting propagation and basal cut configuration

A 2023 study examined hardwood cuttings of *Cydonia oblonga*, *Punica granatum*, and *Ficus carica* under two basal cut directions—**straight** and **slant** at **45°**—and five basal internode stub lengths below the basal node: **0 cm**, **0.5 cm**, **1.0 cm**, **2.0 cm**, and **3.0 cm**. A key design detail was that **slant cuts were not tested at 0 cm**. The experiment used a **RCBD with three replications**, **6 cuttings per bag**, and **486 cuttings** in total. Cuttings were **20 cm** long, **0.7–1.1 mm** in diameter, taken from basal parts of one-year-old shoots, treated with captan, rooted in sand, and evaluated for **rooting percentage, root number, root length, shoot length, shoot diameter, and survival percentage** [2311.04953].

For *Punica granatum*, **basal cut direction alone** did not significantly affect rooting percentage or the other measured traits. The reported values were **37.77%** rooting for **straight** cuts and **33.33%** for **slant** cuts, with **100%** survival in both cases. Root number was **8.12** for straight cuts and **6.77** for slant cuts; root length was **2.12 cm** and **1.67 cm**, respectively; shoot length was **8.72 cm** and **7.03 cm**; shoot diameter was **2.08 mm** and **1.91 mm**. All differences were reported as not significant at \(P \le 0.05\).

By contrast, **basal internode stub length alone** significantly affected pomegranate rooting percentage. The reported rooting percentages were **33.33%** at **0.0 cm**, **22.21%** at **0.5 cm**, **44.44%** at **1.0 cm**, **44.44%** at **2.0 cm**, and **33.33%** at **3.0 cm**. The highest rooting percentage was therefore **44.44%**, obtained with **1 and 2 cm basal internode stub lengths**, whereas **0.5 cm** produced the lowest rooting percentage, **22.21%**. Root number also varied significantly, reaching **12.55** at **2.0 cm** and falling to **2.96** at **3.0 cm**. Root length, shoot length, shoot diameter, and survival were not significantly affected by stub length, and survival remained **100%** across all pomegranate treatments.

The decisive result came from the **interaction** between basal cut direction and stub length. For pomegranate, this interaction was significant for rooting percentage and several other traits. The **best rooting capacity** was **49.99%**, achieved when cuttings were **straightly cut at the base with 1 cm and 2 cm basal internode stub lengths**. The **lowest rooting percentage** was **16.66%** for **slant cut with 0.5 cm basal internode stub length**. In the interaction table, straight cuts gave **33.33%** at **0.0 cm**, **27.77%** at **0.5 cm**, **49.99%** at **1.0 cm**, **49.99%** at **2.0 cm**, and **27.77%** at **3.0 cm**; slant cuts gave **16.66%** at **0.5 cm** and **38.88%** at **1.0 cm**, **2.0 cm**, and **3.0 cm**. The highest root number, **13.88**, occurred in **straight cut + 2 cm stub**, whereas the lowest, **2.70**, occurred in **slant cut + 3 cm stub**. A common practical assumption that slant cuts improve rooting is not supported for pomegranate in this study; the reported best combinations were straight, not slant.

## 3. Genotype-dependent rooting and irrigation frequency

A 2024 study investigated rooting of pomegranate hardwood cuttings from **11 genotypes** under four irrigation frequencies in an **uncontrolled greenhouse**: **1-day**, **2-day**, **7-day**, and **10-day** intervals. Hardwood cuttings were taken from the **basal part of one-year-old shoots** on **5 February 2022**. Each cutting was **15 ± 1 cm** long and **0.6–1 cm** in diameter. The genotypes were **G1 — Salakhani trsh**, **G2 — Salakhani mekhosh**, **G3 — Amriki**, **G4 — Twekl sury trsh**, **G5 — Twekl astury naw spy**, **G6 — Hanara sherina**, **G7 — Kawa hanary sherin**, **G8 — Kawa hanary trsh**, **G9 — Malesay twekl asture**, **G10 — Malesay twekl tank**, and **G11 — Sura hanary trsh**. The design was a **Randomized Complete Block Design (RCBD)** with **60 cuttings per genotype**, divided into four irrigation groups of **15 cuttings per genotype**, arranged as **3 replications of 5 cuttings each**. Cuttings were planted in **sand** in polyethylene bags and watered by hand. Greenhouse conditions ranged from **14.7°C to 28.2°C** and **37% to 62% relative humidity**. After **15 weeks**, measurements included rooting percentage, root number, longest root length, longest shoot length, shoot diameter, leaf number, leaf area, and relative chlorophyll content (SPAD); the analysis used **ANOVA**, **Duncan’s Multiple Range Test**, **Pearson correlation**, **PCA**, and **UPGMA cluster analysis** [2407.00408].

Genotype had a strong and significant effect on rooting. The highest rooting percentages were **G11 = 95%**, **G6 = 90%**, and **G7 = 83%**. Intermediate values were **G8 = 71%**, **G2 = 63%**, **G4 = 63%**, and **G9 = 63%**. The poorest rooting genotypes were **G1 = 28%**, **G5 = 36%**, **G3 = 38%**, and **G10 = 40%**. Table 2 further reported **G11 = 95 a**, **G6 = 90 a**, and **G7 = 83 a**, supporting the conclusion that these genotypes had the strongest rooting ability. The same table reported that **G6** had the **highest root number (54.62)** and **longest roots (16.77 cm)**, **G4** had **shoot length 20.00 cm** and root number **52.60**, and **G11** had the **highest leaf number (27.54)**.

The irrigation result was genotype-dependent, but the paper’s practical conclusion was unequivocal: **7-day irrigation was the best overall frequency**. The abstract states that a **7-day frequency** produced the maximum rooting percentages in **G6 (93)**, **G9 (86)**, **G2 (80)**, **G4 (73)**, **G3 (53)**, and **G1 (40)**. The minimum rooting percentage, **20%**, occurred in **G3 with a 1-day frequency** and in **G1 with 10-day frequency**. For **G5, G7, G8, G10, and G11**, rooting percentage was **not significantly affected** by irrigation frequency. The study therefore concluded that “**irrigation with a 7-day frequency was the best for the cuttings of all the pomegranate genotypes investigated**.”

The multivariate analyses reinforced this interpretation. Pearson analysis found significant positive relationships between **root number and shoot length** (**r = 0.68, P = 0.02**), **root number and leaf area** (**r = 0.67, P = 0.02**), **root length and shoot length** (**r = 0.65, P = 0.03**), and **root length and leaf area** (**r = 0.68, P = 0.02**). **PCA** explained **67.73%** of the total variance, with **PCA1 = 45.03%** and **PCA2 = 22.7%**. In that space, **G6 and G4** were associated with **root number, root length, and leaf area**, whereas **G11** was associated with **rooting percentage, leaf number, shoot length, and shoot diameter**. The **UPGMA cluster analysis** separated the genotypes into **three major clusters**, with **G6, G7, and G11** forming **Cluster 3**, distinguished by the best cutting performance overall. This suggests that pomegranate rooting in nursery propagation is not only irrigation-sensitive but strongly structured by genotype.

## 4. Single-beam rotational imaging of *Punica granatum* pollen

In a 2025 optics study, *Punica granatum* pollen cells served as one of the principal biological validation samples for a **single-beam driven rotational manipulation** method designed for **high-resolution 3D cellular morphology reconstruction**. The method uses a **single tightly focused optical beam** carrying **spin angular momentum (SAM)** to **trap** and **rotate** cells while a camera remains **coaxially aligned with the optical axis**, allowing acquisition of multi-view images without moving the camera or sample stage. The study explicitly motivates this design by the need to overcome the **missing-cone problem** and the degraded axial resolution associated with single-view imaging [2506.07145].

The paper expresses SAM density for a monochromatic field as
\[
\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],
\]
and writes the circularly polarized plane wave as
\[
\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.
\]
Its central manipulation concept is not direct circular polarization, but a **linearly polarized beam** incident at **45°** relative to the SLM-controlled polarization axis, combined with a **polarization-sensitive SLM** that generates a focused field with a **tunable SAM direction**. The resulting SAM vector is written as
\[
\mathbf{S}=S
\begin{bmatrix}
\sin\Theta\cos\Phi\,\hat{\mathbf e}_x\\[4pt]
\sin\Theta\sin\Phi\,\hat{\mathbf e}_y\\[4pt]
\cos\Theta\,\hat{\mathbf e}_z
\end{bmatrix}.
\]
By choosing \((\Theta,\Phi)\), the system can rotate a trapped cell about arbitrary axes, including transverse axes relevant for multi-view acquisition.

The experimental setup was a holographic optical tweezers system based on an inverted microscope: **Olympus IX73**, **532 nm continuous-wave solid-state laser**, **4f system** with lenses \(L_1\) and \(L_2\), **HOLOEYE Pluto** SLM with **\(1920 \times 1080\)** pixels, **half-wave plate**, **oil-immersion 10× objective with NA = 1.3**, and **CMOS camera (PixelLINK PL-D752MU)**. The reconstruction workflow comprised six steps: rotation and image acquisition; cell detection and cropping using **YOLOv8** with training data annotated in **LabelImg**; segmentation using a **Yolov11-based segmentation approach** with masks prepared in **LabelMe**; orientation and position normalization; **visual hull** reconstruction using
\[
V_{\text{3D}}=\bigcap_{i=1}^{N} V_i;
\]
and rendering plus quantitative analysis. Spatial calibration was **1 pixel = 0.1372 μm**, giving an isotropic voxel size of
\[
0.1372 \times 0.1372 \times 0.1372\ \mu\text{m}^3.
\]
Volume was computed by voxel counting,
\[
V = N_{\text{vox}}\cdot(0.1372)^3\ \mu\text{m}^3.
\]

The reconstructed *Punica granatum* pollen cell was reported as **roughly spherical** and **symmetric**, with **three prominent protrusions**. The color-mapped distance-to-centroid visualization showed the strongest depth variations at those three apical regions, whereas most of the remaining surface showed relatively small deviation from the centroid. The reported volume was **6021.3 μm³**. By comparison, the reconstructed *Prunus cerasifera* cell was **ellipsoidal and irregular**, with a volume of **1173.1 μm³**. The pomegranate result therefore demonstrated a compact, nearly spherical body with distinct apex-like features that would be difficult to infer from a single 2D projection.

## 5. Punica as a multi-tenant LoRA serving architecture

In machine-learning systems research, “Punica” denotes a serving system for **multiple LoRA models in a shared GPU cluster**. The motivation arises from the structure of LoRA itself: a pretrained weight matrix \(W\in\mathbb{R}^{h_1\times h_2}\) is adapted to
\[
W + AB,
\]
where \(A\in\mathbb{R}^{h_1\times r}\), \(B\in\mathbb{R}^{r\times h_2}\), and \(r\ll h_1,h_2\). The paper states that each adapted model only adds about **0.1% to 1%** of the base model weight. Punica is built around three guidelines: **consolidate workloads onto as few GPUs as possible**, **enable batching across different LoRA models**, and **focus on the decode stage**, which dominates serving cost for long generations [2310.18547].

The system architecture includes frontends exposing a **REST API**, a **centralized scheduler**, and per-GPU **runners**. Each user request contains a prompt and a **LoRA model ID**. Each GPU loads the full pretrained backbone **once**, a large **KvCache** reservation, and only the LoRA adapter weights \(A,B\) for the tenants active on that GPU. The paper reports LoRA loading costs of around **50 μs per layer** and about **2 ms for the whole model** on **PCIe Gen4 x16**, with the load overlapping ongoing GPU computation. This architecture operationalizes the shared-backbone premise: the base model is not replicated per tenant.

Punica’s main kernel contribution is **Segmented Gather Matrix-Vector multiplication (SGMV)**, which enables batched LoRA computation even when each request in a batch uses a different LoRA model. The LoRA add-on \(\mathbf{x}AB\) is decomposed into two launches: **SGMV-shrink** for \(\mathbf{v}=\mathbf{x}A\) and **SGMV-expand** for \(\mathbf{y}=\mathbf{v}B\). The kernel scheduling differs between these cases. For **expand**, the output dimension is large enough to expose parallelism, so the output matrix is split across thread blocks. For **shrink**, the output is too thin, so Punica uses a **Split-K** strategy that splits the input dimension across blocks and then reduces partial results. In both cases, the LoRA model index is bound to `blockIdx.y`. For the special case in which every request has a distinct LoRA model, the operator becomes effectively matrix-vector multiplication and is treated as **memory-bandwidth bound**, with a specialized schedule that avoids Tensor Cores.

The paper contrasts SGMV with a **gather-and-bmm** implementation that first stacks weights and then calls `torch.bmm()`. It attributes the disadvantage of Gather-BMM to extra memory traffic, specifically an additional
\[
s_n \times h_i \times h_o \times 2
\]
elements of memory I/O compared with SGMV. The roofline analysis is written as
\[
\mathrm{FLOP} = s_n \times h_i \times h_o \times 2
\]
and
\[
\mathrm{I/O} = [s_n \times (h_i+h_o) + n \times h_1 \times h_2] \times 2.
\]
The scheduler is likewise specialized for shared-cluster serving. For each GPU it tracks the current working set or batch size, available KvCache memory, and whether the GPU has reached the batch-size cap. A new request is assigned to the GPU with the **largest working set** that still satisfies memory and max-batch constraints; ties are broken by the **highest GPU UUID**; if all GPUs are saturated, the request is queued **FCFS**. On **A100 GPUs**, the paper reports **32** as a good max batch size. Migration is implemented by **evict/cancel** on the source GPU and **re-add** on another GPU, with recomputation of the prefix rather than copying the KvCache.

## 6. Evaluation, batching semantics, and deployment implications

Punica also changes the KvCache layout to support continuous batching through a paged, separable organization:
\[
\left[\sum_i \left\lceil \frac{S_i}{P} \right\rceil, L, 2, N, P, D\right],
\]
where \(S_i\) is sequence length for request \(i\), \(P\) is page size, \(L\) is the number of layers, \(N\) is the number of heads, and \(D\) is head dimension. The implementation is a **PyTorch extension** exposing custom CUDA kernels via **PyBind11**, with a Python runtime adapted from **HuggingFace Transformers**, **FlashInfer** for efficient attention and paged cache support, **fused LayerNorm**, and **Rust-based scheduler/frontend/runner components**. A notable detail is that Punica mixes **prefill and decode requests in the same batch** where possible, with prefill requests at the beginning and decode requests later, using a `BatchLen` structure whose segmentation and SGMV indices are built once per model invocation and reused across all layers [2310.18547].

The evaluation used **one NVIDIA A100 80GB GPU** for microbenchmarks and single-GPU serving, and **two HGX A100 40GB servers with 8 GPUs each** for **70B tensor-parallel experiments** and cluster deployment. Models were **Llama-2 7B, 13B, and 70B**; LoRA rank was **16** and applied to all dense projections. Workloads followed ShareGPT prompt/response length distributions with four LoRA-popularity patterns: **Distinct**, **Uniform**, **Skewed**, and **Identical**. Baselines were **HuggingFace Transformers + PEFT**, **DeepSpeed Inference**, **FasterTransformer**, and **vLLM**.

The performance claims are explicit. On a single A100, Punica achieved **1044 tok/s** on **Llama-2 7B** and **693 tok/s** on **Llama-2 13B**. In the **70B** tensor-parallel setting on **8 GPUs**, Punica reached about **441–446 tok/s**, whereas **vLLM** dropped to around **21–25 tok/s** on multiple LoRA models. In the cluster experiment with **16 GPUs**, Punica consolidated requests onto fewer GPUs while maintaining high throughput; busy GPUs ran near the maximum batch size, some GPUs became idle and remained idle, and requests were migrated when KvCache pressure rose. The paper’s headline system-level result is that, with a fixed-sized GPU cluster, Punica achieved **12x higher throughput** in serving multiple LoRA models compared to state-of-the-art LLM serving systems while adding only **2 ms latency per token**.

The paper also states the system’s tradeoffs and limitations. It focuses primarily on the **decode stage**; it uses **recomputation for migration** instead of copying KvCache; it requires a scheduler that tracks GPU memory and batch occupancy carefully; and it is most effective when there is enough concurrent demand to form large batches. Workloads with extremely high tenant churn may still incur adapter-load overheads, though the paper characterizes this as millisecond-scale. In this sense, Punica is not merely a kernel optimization: it is a full serving stack whose performance depends jointly on segmented LoRA execution, continuous batching, KvCache layout, and consolidation-aware scheduling.

Source: https://www.emergentmind.com/topics/punica