Papers
Topics
Authors
Recent
Search
2000 character limit reached

Punica: Pomegranate Propagation & LLM Serving

Updated 18 July 2026
  • Punica is a term with dual definitions, representing both botanical studies on pomegranate propagation and microscopic pollen morphology, and a high-performance, multi-tenant serving system for LoRA models.
  • In propagation research, straight basal cuts with 1–2 cm internode stubs resulted in optimal rooting performance, achieving up to 49.99% rooting in pomegranate cuttings.
  • In systems research, Punica’s framework delivers 12x higher throughput and only 2 ms latency per token by leveraging a shared pretrained backbone, novel CUDA kernels, and consolidation-aware scheduling.

Punica appears in the cited literature in two distinct technical senses. In plant science and microscopy, it denotes Punica granatum (pomegranate), examined both as hardwood cuttings in propagation experiments and as pollen cells in high-resolution 3D morphology reconstruction. In systems research, “Punica” denotes a multi-tenant LoRA serving system for LLMs that exploits a shared pretrained backbone, a new CUDA kernel, and a consolidation-oriented scheduler for efficient GPU-cluster operation (Mohammed et al., 2023, Salih et al., 2024, Pan et al., 8 Jun 2025, Chen et al., 2023).

1. Scope and research usages

The current literature represented here uses the name across biological and computational domains rather than in a single uniform sense. In the biological papers, the emphasis is on vegetative propagation and cellular morphology of Punica granatum. In the systems paper, the same name identifies an LLM-serving framework for multi-tenant LoRA workloads.

Usage of “Punica” Object studied Representative result
Punica granatum in propagation Hardwood cuttings Best interaction treatment: 49.99% rooting with straight basal cut + 1 cm or 2 cm basal internode stub
Punica granatum in microscopy Pollen cell Reconstructed as roughly spherical with three prominent protrusions; volume 6021.3 μm³
Punica as a systems framework Multi-tenant LoRA serving 12x higher throughput with only 2 ms latency per token

This multiplicity of usage is central to interpreting the term in recent technical writing: in one branch it functions as a biological specimen designation, and in another as a proper name for an inference-serving system.

2. Hardwood cutting propagation and basal cut configuration

A 2023 study examined hardwood cuttings of Cydonia oblonga, Punica granatum, and Ficus carica under two basal cut directions—straight and slant at 45°—and five basal internode stub lengths below the basal node: 0 cm, 0.5 cm, 1.0 cm, 2.0 cm, and 3.0 cm. A key design detail was that slant cuts were not tested at 0 cm. The experiment used a RCBD with three replications, 6 cuttings per bag, and 486 cuttings in total. Cuttings were 20 cm long, 0.7–1.1 mm in diameter, taken from basal parts of one-year-old shoots, treated with captan, rooted in sand, and evaluated for rooting percentage, root number, root length, shoot length, shoot diameter, and survival percentage (Mohammed et al., 2023).

For Punica granatum, basal cut direction alone did not significantly affect rooting percentage or the other measured traits. The reported values were 37.77% rooting for straight cuts and 33.33% for slant cuts, with 100% survival in both cases. Root number was 8.12 for straight cuts and 6.77 for slant cuts; root length was 2.12 cm and 1.67 cm, respectively; shoot length was 8.72 cm and 7.03 cm; shoot diameter was 2.08 mm and 1.91 mm. All differences were reported as not significant at P0.05P \le 0.05.

By contrast, basal internode stub length alone significantly affected pomegranate rooting percentage. The reported rooting percentages were 33.33% at 0.0 cm, 22.21% at 0.5 cm, 44.44% at 1.0 cm, 44.44% at 2.0 cm, and 33.33% at 3.0 cm. The highest rooting percentage was therefore 44.44%, obtained with 1 and 2 cm basal internode stub lengths, whereas 0.5 cm produced the lowest rooting percentage, 22.21%. Root number also varied significantly, reaching 12.55 at 2.0 cm and falling to 2.96 at 3.0 cm. Root length, shoot length, shoot diameter, and survival were not significantly affected by stub length, and survival remained 100% across all pomegranate treatments.

The decisive result came from the interaction between basal cut direction and stub length. For pomegranate, this interaction was significant for rooting percentage and several other traits. The best rooting capacity was 49.99%, achieved when cuttings were straightly cut at the base with 1 cm and 2 cm basal internode stub lengths. The lowest rooting percentage was 16.66% for slant cut with 0.5 cm basal internode stub length. In the interaction table, straight cuts gave 33.33% at 0.0 cm, 27.77% at 0.5 cm, 49.99% at 1.0 cm, 49.99% at 2.0 cm, and 27.77% at 3.0 cm; slant cuts gave 16.66% at 0.5 cm and 38.88% at 1.0 cm, 2.0 cm, and 3.0 cm. The highest root number, 13.88, occurred in straight cut + 2 cm stub, whereas the lowest, 2.70, occurred in slant cut + 3 cm stub. A common practical assumption that slant cuts improve rooting is not supported for pomegranate in this study; the reported best combinations were straight, not slant.

3. Genotype-dependent rooting and irrigation frequency

A 2024 study investigated rooting of pomegranate hardwood cuttings from 11 genotypes under four irrigation frequencies in an uncontrolled greenhouse: 1-day, 2-day, 7-day, and 10-day intervals. Hardwood cuttings were taken from the basal part of one-year-old shoots on 5 February 2022. Each cutting was 15 ± 1 cm long and 0.6–1 cm in diameter. The genotypes were G1 — Salakhani trsh, G2 — Salakhani mekhosh, G3 — Amriki, G4 — Twekl sury trsh, G5 — Twekl astury naw spy, G6 — Hanara sherina, G7 — Kawa hanary sherin, G8 — Kawa hanary trsh, G9 — Malesay twekl asture, G10 — Malesay twekl tank, and G11 — Sura hanary trsh. The design was a Randomized Complete Block Design (RCBD) with 60 cuttings per genotype, divided into four irrigation groups of 15 cuttings per genotype, arranged as 3 replications of 5 cuttings each. Cuttings were planted in sand in polyethylene bags and watered by hand. Greenhouse conditions ranged from 14.7°C to 28.2°C and 37% to 62% relative humidity. After 15 weeks, measurements included rooting percentage, root number, longest root length, longest shoot length, shoot diameter, leaf number, leaf area, and relative chlorophyll content (SPAD); the analysis used ANOVA, Duncan’s Multiple Range Test, Pearson correlation, PCA, and UPGMA cluster analysis (Salih et al., 2024).

Genotype had a strong and significant effect on rooting. The highest rooting percentages were G11 = 95%, G6 = 90%, and G7 = 83%. Intermediate values were G8 = 71%, G2 = 63%, G4 = 63%, and G9 = 63%. The poorest rooting genotypes were G1 = 28%, G5 = 36%, G3 = 38%, and G10 = 40%. Table 2 further reported G11 = 95 a, G6 = 90 a, and G7 = 83 a, supporting the conclusion that these genotypes had the strongest rooting ability. The same table reported that G6 had the highest root number (54.62) and longest roots (16.77 cm), G4 had shoot length 20.00 cm and root number 52.60, and G11 had the highest leaf number (27.54).

The irrigation result was genotype-dependent, but the paper’s practical conclusion was unequivocal: 7-day irrigation was the best overall frequency. The abstract states that a 7-day frequency produced the maximum rooting percentages in G6 (93), G9 (86), G2 (80), G4 (73), G3 (53), and G1 (40). The minimum rooting percentage, 20%, occurred in G3 with a 1-day frequency and in G1 with 10-day frequency. For G5, G7, G8, G10, and G11, rooting percentage was not significantly affected by irrigation frequency. The study therefore concluded that “irrigation with a 7-day frequency was the best for the cuttings of all the pomegranate genotypes investigated.”

The multivariate analyses reinforced this interpretation. Pearson analysis found significant positive relationships between root number and shoot length (r = 0.68, P = 0.02), root number and leaf area (r = 0.67, P = 0.02), root length and shoot length (r = 0.65, P = 0.03), and root length and leaf area (r = 0.68, P = 0.02). PCA explained 67.73% of the total variance, with PCA1 = 45.03% and PCA2 = 22.7%. In that space, G6 and G4 were associated with root number, root length, and leaf area, whereas G11 was associated with rooting percentage, leaf number, shoot length, and shoot diameter. The UPGMA cluster analysis separated the genotypes into three major clusters, with G6, G7, and G11 forming Cluster 3, distinguished by the best cutting performance overall. This suggests that pomegranate rooting in nursery propagation is not only irrigation-sensitive but strongly structured by genotype.

4. Single-beam rotational imaging of Punica granatum pollen

In a 2025 optics study, Punica granatum pollen cells served as one of the principal biological validation samples for a single-beam driven rotational manipulation method designed for high-resolution 3D cellular morphology reconstruction. The method uses a single tightly focused optical beam carrying spin angular momentum (SAM) to trap and rotate cells while a camera remains coaxially aligned with the optical axis, allowing acquisition of multi-view images without moving the camera or sample stage. The study explicitly motivates this design by the need to overcome the missing-cone problem and the degraded axial resolution associated with single-view imaging (Pan et al., 8 Jun 2025).

The paper expresses SAM density for a monochromatic field as

S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],

and writes the circularly polarized plane wave as

E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.

Its central manipulation concept is not direct circular polarization, but a linearly polarized beam incident at 45° relative to the SLM-controlled polarization axis, combined with a polarization-sensitive SLM that generates a focused field with a tunable SAM direction. The resulting SAM vector is written as

$\mathbf{S}=S \begin{bmatrix} \sin\Theta\cos\Phi\,\hat{\mathbf e}_x\[4pt] \sin\Theta\sin\Phi\,\hat{\mathbf e}_y\[4pt] \cos\Theta\,\hat{\mathbf e}_z \end{bmatrix}.$

By choosing (Θ,Φ)(\Theta,\Phi), the system can rotate a trapped cell about arbitrary axes, including transverse axes relevant for multi-view acquisition.

The experimental setup was a holographic optical tweezers system based on an inverted microscope: Olympus IX73, 532 nm continuous-wave solid-state laser, 4f system with lenses L1L_1 and L2L_2, HOLOEYE Pluto SLM with 1920×10801920 \times 1080 pixels, half-wave plate, oil-immersion 10× objective with NA = 1.3, and CMOS camera (PixelLINK PL-D752MU). The reconstruction workflow comprised six steps: rotation and image acquisition; cell detection and cropping using YOLOv8 with training data annotated in LabelImg; segmentation using a Yolov11-based segmentation approach with masks prepared in LabelMe; orientation and position normalization; visual hull reconstruction using

V3D=i=1NVi;V_{\text{3D}}=\bigcap_{i=1}^{N} V_i;

and rendering plus quantitative analysis. Spatial calibration was 1 pixel = 0.1372 μm, giving an isotropic voxel size of

0.1372×0.1372×0.1372 μm3.0.1372 \times 0.1372 \times 0.1372\ \mu\text{m}^3.

Volume was computed by voxel counting,

S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],0

The reconstructed Punica granatum pollen cell was reported as roughly spherical and symmetric, with three prominent protrusions. The color-mapped distance-to-centroid visualization showed the strongest depth variations at those three apical regions, whereas most of the remaining surface showed relatively small deviation from the centroid. The reported volume was 6021.3 μm³. By comparison, the reconstructed Prunus cerasifera cell was ellipsoidal and irregular, with a volume of 1173.1 μm³. The pomegranate result therefore demonstrated a compact, nearly spherical body with distinct apex-like features that would be difficult to infer from a single 2D projection.

5. Punica as a multi-tenant LoRA serving architecture

In machine-learning systems research, “Punica” denotes a serving system for multiple LoRA models in a shared GPU cluster. The motivation arises from the structure of LoRA itself: a pretrained weight matrix S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],1 is adapted to

S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],2

where S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],3, S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],4, and S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],5. The paper states that each adapted model only adds about 0.1% to 1% of the base model weight. Punica is built around three guidelines: consolidate workloads onto as few GPUs as possible, enable batching across different LoRA models, and focus on the decode stage, which dominates serving cost for long generations (Chen et al., 2023).

The system architecture includes frontends exposing a REST API, a centralized scheduler, and per-GPU runners. Each user request contains a prompt and a LoRA model ID. Each GPU loads the full pretrained backbone once, a large KvCache reservation, and only the LoRA adapter weights S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],6 for the tenants active on that GPU. The paper reports LoRA loading costs of around 50 μs per layer and about 2 ms for the whole model on PCIe Gen4 x16, with the load overlapping ongoing GPU computation. This architecture operationalizes the shared-backbone premise: the base model is not replicated per tenant.

Punica’s main kernel contribution is Segmented Gather Matrix-Vector multiplication (SGMV), which enables batched LoRA computation even when each request in a batch uses a different LoRA model. The LoRA add-on S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],7 is decomposed into two launches: SGMV-shrink for S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],8 and SGMV-expand for S=12ω0Im[ε0E×E],\mathbf{S}=\frac{1}{2\omega_0}\operatorname{Im}\left[\varepsilon_0 \mathbf{E}\times \mathbf{E}^{*}\right],9. The kernel scheduling differs between these cases. For expand, the output dimension is large enough to expose parallelism, so the output matrix is split across thread blocks. For shrink, the output is too thin, so Punica uses a Split-K strategy that splits the input dimension across blocks and then reduces partial results. In both cases, the LoRA model index is bound to blockIdx.y. For the special case in which every request has a distinct LoRA model, the operator becomes effectively matrix-vector multiplication and is treated as memory-bandwidth bound, with a specialized schedule that avoids Tensor Cores.

The paper contrasts SGMV with a gather-and-bmm implementation that first stacks weights and then calls torch.bmm(). It attributes the disadvantage of Gather-BMM to extra memory traffic, specifically an additional

E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.0

elements of memory I/O compared with SGMV. The roofline analysis is written as

E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.1

and

E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.2

The scheduler is likewise specialized for shared-cluster serving. For each GPU it tracks the current working set or batch size, available KvCache memory, and whether the GPU has reached the batch-size cap. A new request is assigned to the GPU with the largest working set that still satisfies memory and max-batch constraints; ties are broken by the highest GPU UUID; if all GPUs are saturated, the request is queued FCFS. On A100 GPUs, the paper reports 32 as a good max batch size. Migration is implemented by evict/cancel on the source GPU and re-add on another GPU, with recomputation of the prefix rather than copying the KvCache.

6. Evaluation, batching semantics, and deployment implications

Punica also changes the KvCache layout to support continuous batching through a paged, separable organization: E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.3 where E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.4 is sequence length for request E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.5, E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.6 is page size, E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.7 is the number of layers, E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.8 is the number of heads, and E=e^x+iσe^y2E0eikxiωt.\mathbf{E}=\frac{\hat{\mathbf e}_x+i\sigma \hat{\mathbf e}_y}{2}\,E_0\,e^{i\mathbf{k}\cdot \mathbf{x}-i\omega t}.9 is head dimension. The implementation is a PyTorch extension exposing custom CUDA kernels via PyBind11, with a Python runtime adapted from HuggingFace Transformers, FlashInfer for efficient attention and paged cache support, fused LayerNorm, and Rust-based scheduler/frontend/runner components. A notable detail is that Punica mixes prefill and decode requests in the same batch where possible, with prefill requests at the beginning and decode requests later, using a BatchLen structure whose segmentation and SGMV indices are built once per model invocation and reused across all layers (Chen et al., 2023).

The evaluation used one NVIDIA A100 80GB GPU for microbenchmarks and single-GPU serving, and two HGX A100 40GB servers with 8 GPUs each for 70B tensor-parallel experiments and cluster deployment. Models were Llama-2 7B, 13B, and 70B; LoRA rank was 16 and applied to all dense projections. Workloads followed ShareGPT prompt/response length distributions with four LoRA-popularity patterns: Distinct, Uniform, Skewed, and Identical. Baselines were HuggingFace Transformers + PEFT, DeepSpeed Inference, FasterTransformer, and vLLM.

The performance claims are explicit. On a single A100, Punica achieved 1044 tok/s on Llama-2 7B and 693 tok/s on Llama-2 13B. In the 70B tensor-parallel setting on 8 GPUs, Punica reached about 441–446 tok/s, whereas vLLM dropped to around 21–25 tok/s on multiple LoRA models. In the cluster experiment with 16 GPUs, Punica consolidated requests onto fewer GPUs while maintaining high throughput; busy GPUs ran near the maximum batch size, some GPUs became idle and remained idle, and requests were migrated when KvCache pressure rose. The paper’s headline system-level result is that, with a fixed-sized GPU cluster, Punica achieved 12x higher throughput in serving multiple LoRA models compared to state-of-the-art LLM serving systems while adding only 2 ms latency per token.

The paper also states the system’s tradeoffs and limitations. It focuses primarily on the decode stage; it uses recomputation for migration instead of copying KvCache; it requires a scheduler that tracks GPU memory and batch occupancy carefully; and it is most effective when there is enough concurrent demand to form large batches. Workloads with extremely high tenant churn may still incur adapter-load overheads, though the paper characterizes this as millisecond-scale. In this sense, Punica is not merely a kernel optimization: it is a full serving stack whose performance depends jointly on segmented LoRA execution, continuous batching, KvCache layout, and consolidation-aware scheduling.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Punica.