---
title: 'MVR-cache: Multi-Vector LLM Caching'
url: https://www.emergentmind.com/topics/mvr-cache
type: topic
---

# MVR-cache: Multi-Vector LLM Caching

MVR-cache is a semantic caching system for LLMs that replaces single-vector cosine similarity with a learned, segment-aware multi-vector similarity. In the formulation introduced in “MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation” [2605.24914], each prompt is segmented by a learned policy, embedded as multiple segment vectors, compared to cached prompts through a symmetric normalized MaxSim score, and then filtered through a vCache-style probabilistic correctness model. The stated objective is to improve cache hit rate without relaxing correctness guarantees, and the system is trained with a reinforcement-learning procedure derived from a theoretical analysis of the similarity–correctness relationship [2605.24914].

## 1. Definition and problem setting

MVR-cache addresses semantic caching for LLM APIs, where a system must decide whether a new prompt \(x\) is sufficiently close to a cached prompt \(x_i\) that it can safely reuse the cached response \(r(x_i)\) instead of issuing a new LLM call [2605.24914]. In this setting, existing systems are described as representing each prompt with a single embedding vector \(\phi(x) \in \mathbb{R}^d\), using a scalar similarity such as cosine similarity, and then applying thresholds to decide reuse [2605.24914].

The central claim of MVR-cache is that this single-vector formulation is too coarse for prompts with multiple semantic aspects. The paper states that a single vector cannot capture multiple distinct semantic aspects of a longer or complex prompt, and that cosine similarity often ranks prompts that are topically similar but response-different above prompts that are actually response-equivalent [2605.24914]. MVR-cache therefore shifts the design focus from threshold tuning to the similarity layer itself.

Within this framework, semantic caching matters because a cache hit avoids an LLM call, and each LLM call is characterized as orders of magnitude more expensive and slower than embedding plus retrieval. The paper treats correctness as non-negotiable: it keeps the same correctness-guarantee framework as vCache, while attempting to make nearest-neighbor retrieval more faithful to response equivalence [2605.24914].

A plausible implication is that MVR-cache belongs to a broader class of retrieval-augmented cache controllers in which the main bottleneck is not storage but semantic matching fidelity. In that sense, its main contribution is not a new cache eviction policy, but a new similarity representation and training procedure for hit-or-miss decisions.

## 2. Representation: learned segmentation and multi-vector retrieval

The system begins by identifying candidate split positions \(P_x = \{p_1,\dots,p_{|P_x|}\}\) for a prompt \(x\), such as token indices of punctuation marks [2605.24914]. A learned segmentation model \(\mathcal{E}_\theta\) then chooses an ordered subset of these positions,
\[
\mathcal{E}_\theta(x)=p=[p_{i_1},\dots,p_{i_{m-1}}],
\]
which induces contiguous segments \(x^{(1)},\dots,x^{(m)}\) [2605.24914]. Each segment is embedded with a shared encoder \(\text{Emb}(\cdot)\), yielding a multi-vector representation rather than a single prompt embedding.

The segmentation model is implemented as a pointer network. The prompt tokens are encoded by BERT, projected by an MLP, and processed by a single-layer LSTM decoder. At each decoding step, attention scores
\[
u_{k,j}=v^\top \tanh(W_1\mathbf{h}_j + W_2\mathbf{d}_k)
\]
are computed over candidate split positions, masked to enforce increasing boundary order, and the process continues until a special <stop> token is chosen [2605.24914]. The resulting segmentation policy is denoted \(\pi_\theta(p\mid x)\equiv T_\theta(p\mid x)\).

Once segmented, a prompt is represented as
\[
\big(\mathbf{e}(x^{(1)}),\dots,\mathbf{e}(x^{(m)})\big)\in\mathbb{R}^{m\times d}.
\]
Similarity between a query prompt and a cached prompt is then computed with ColBERT-style MaxSim. For prompts \(x\) and \(x_j\), unidirectional MaxSim is
\[
\text{MaxSim}(x,x_j)=\sum_{t=1}^{m}\max_{s=1}^{m_j}\text{sim}\big(\mathbf{e}(x^{(t)}),\mathbf{e}(x_j^{(s)})\big),
\]
where \(\text{sim}(\cdot,\cdot)\) is a base similarity such as cosine similarity [2605.24914].

For semantic caching, the paper defines a symmetric normalized score,
\[
\text{SMaxSim}_\theta(x_i,x_j)
=
\tfrac{1}{2}
\left(
\frac{1}{|x_i|}\text{MaxSim}(x_i,x_j)
+
\frac{1}{|x_j|}\text{MaxSim}(x_j,x_i)
\right),
\]
and retrieves the nearest neighbor by
\[
\text{nn}_\theta(x)=\arg\max_j \text{SMaxSim}_\theta(x,x_j),\qquad
s_\theta(x)=\text{SMaxSim}_\theta(x,\text{nn}_\theta(x)).
\]
The normalization by segment count is intended to make scores comparable across prompts of different lengths, and the symmetrization is intended to enforce mutual relevance rather than one-sided containment [2605.24914].

This design distinguishes MVR-cache from single-vector semantic caches. The paper’s argument is that partial matching at segment level can preserve response-determining information that would be flattened away by a whole-prompt embedding. A plausible implication is that MVR-cache is best viewed as a retrieval model specialized for cacheability rather than as a general-purpose semantic search model.

## 3. Correctness model and theoretical objective

MVR-cache preserves the vCache-style correctness model. For a prompt \(x\), with nearest neighbor \(\text{nn}_\theta(x)\) and similarity score \(s_\theta(x)\), the probability that reuse is correct is modeled as
\[
\Pr(c(x)=1\mid s_\theta(x))
=
\sigma_\gamma(s_\theta(x);t)
=
\frac{1}{1+e^{-\gamma(s_\theta(x)-t)}}.
\]
Here \(c(x)\in\{0,1\}\) indicates whether the new response equals the cached response, and \(t,\gamma\) are estimated by maximum likelihood from neighbor-specific metadata \(\mathcal{O}(x_i)\) [2605.24914].

For each cached prompt \(x_i\), the paper defines
\[
\mathcal{O}(x_i)=\{(s(x_j),c(x_j))\mid \text{nn}(x_j)=x_i,\ j>i\},
\]
and fits
\[
(t_i,\gamma_i)
=
\arg\min_{t,\gamma}
\sum_{(s_j,c_j)\in\mathcal{O}(x_i)}
\text{BCE}\big(\sigma_\gamma(s_j;t),c_j\big),
\]
where
\[
\text{BCE}(\hat p,c)=-c\log \hat p -(1-c)\log(1-\hat p).
\]
Given these parameters, vCache computes an exploration probability \(T\) that guarantees overall error \(\le \delta\) under its assumptions [2605.24914].

The theoretical analysis in MVR-cache is built on two assumptions. Assumption 3.1 states that, for fixed \(\theta\), similarity scores conditioned on correctness satisfy
\[
s_\theta(x)\mid c \sim \mathcal{N}(\mu_c,\sigma^2),\qquad \mu_1\neq\mu_0.
\]
Assumption 3.2 states a balanced prior \(\Pr(c=1)=\Pr(c=0)=0.5\) [2605.24914]. Under these assumptions, the paper gives Theorem 3.3: minimizing the MLE/BCE loss with similarity scores \(s_\theta(x)\) is equivalent to maximizing the cache hit rate subject to any user-specified error bound \(\delta\) [2605.24914].

The argument relies on the fact that, under the Gaussian assumption, the posterior correctness probability is exactly logistic with
\[
\gamma=\frac{\mu_1-\mu_0}{\sigma^2},\qquad
t=\frac{\mu_1+\mu_0}{2}.
\]
The population MLE loss depends on the class separation only through \(\Delta=\mu_1-\mu_0\), and minimizing the loss is stated to be equivalent to maximizing \(\Delta\) [2605.24914]. The paper further extends the analysis to imbalanced class priors through a class-rebalanced MLE objective in Lemma 3.4 [2605.24914].

This makes the segmentation model’s training objective unusually explicit. Rather than optimizing a generic ranking loss, MVR-cache optimizes a similarity distribution that is theoretically aligned with the downstream cache-hit objective under the specified correctness constraints. This suggests a direct coupling between retrieval representation learning and cache-control guarantees.

## 4. Reinforcement-learning training procedure

The segmentation policy induces a combinatorial, non-differentiable optimization problem. The chosen split positions affect segment embeddings, segment counts, \(\text{SMaxSim}_\theta\), nearest-neighbor identity, and ultimately the logistic correctness model. Because segmentation decisions are discrete and nearest-neighbor retrieval is itself discrete, the paper formulates training as RL4CO and uses REINFORCE [2605.24914].

For a sampled prompt \(x_i\), the reward is defined as the negative BCE loss accumulated over prompts whose current nearest neighbor is \(x_i\):
\[
R(\theta;x_i)
=
-\sum_{j:\,\text{nn}_\theta(x_j)=x_i}
\text{BCE}\Big(
\sigma_{\gamma_i}(\text{SMaxSim}_\theta(x_i,x_j);t_i),
c_j
\Big).
\]
The expected objective is
\[
\max_\theta
\mathbb{E}_{x_i,p_i,p_j}\big[R(\theta;x_i)\big],
\]
and policy gradients are estimated by standard REINFORCE,
\[
\nabla_\theta J(\theta)
=
\mathbb{E}_{p\sim\pi_\theta}
\big[
R(p)\nabla_\theta \log \pi_\theta(p\mid x)
\big].
\]
The training loop samples prompts, samples segmentations, computes \(\text{SMaxSim}_\theta\), updates \(\theta\), and periodically refits \(t_i,\gamma_i\) by MLE [2605.24914].

A practical complication is that the mapping \(x_j\mapsto \text{nn}_\theta(x_j)\) also depends on \(\theta\). Recomputing all nearest neighbors at every update would be expensive, so the paper freezes the nearest-neighbor mapping for \(K\) steps and then recomputes segmentations and nearest neighbors for the full training set every \(K\) steps [2605.24914]. This decouples frequent policy updates from expensive global reranking.

The following table summarizes the core training components described for MVR-cache.

| Component | Formulation | Role |
|---|---|---|
| Policy | \(\pi_\theta(p\mid x)\) | Samples prompt segmentations |
| Reward | Negative BCE-based loss | Aligns segmentation with correctness-aware hit rate |
| Optimizer | REINFORCE | Handles discrete segmentation actions |

The training procedure is therefore not a post hoc heuristic over an existing embedding model. It is a learned segmentation policy optimized specifically for cache retrieval under a probabilistic correctness model.

## 5. Retrieval pipeline and deployment architecture

At inference time, MVR-cache follows a staged retrieval and decision procedure. Given a prompt \(x\), the system computes candidate positions \(P_x\), segments the prompt with \(\mathcal{E}_\theta\), embeds the segments, retrieves top candidates using a coarse HNSW index over single-vector full-prompt embeddings, reranks the top-20 candidates with \(\text{SMaxSim}_\theta\), and then applies the vCache correctness model to decide exploit versus explore [2605.24914].

The coarse-to-fine retrieval structure is central to the system’s practicality. Each cached prompt stores its segmentation, segment embeddings, response \(r(x_i)\), and metadata \(\mathcal{O}(x_i)\). Retrieval is approximated by a two-stage pipeline: HNSW on full-prompt embeddings narrows the search space, and symmetric MaxSim reranking recovers the more accurate multi-vector similarity needed for semantic caching [2605.24914].

The paper states that the segmentation model uses BERT, projection MLP, single-layer LSTM decoder, and attention; the embedding model is BGE by default; and the LLM used for generating ground-truth training labels is GPT-4o-mini [2605.24914]. Storage overhead consists mainly of multiple vectors per prompt rather than a single vector, but the paper characterizes this overhead as small relative to LLM compute [2605.24914].

The decision rule remains vCache-style. After obtaining the nearest neighbor \(x_i\) and score \(s_\theta(x)\), the system estimates \(\Pr(c(x)=1\mid s_\theta(x))\) from the logistic model associated with \(x_i\), computes an exploration probability \(T\) satisfying the user-provided error constraint \(\delta\), and either returns the cached response \(r(x_i)\) with probability \(1-T\) or calls the LLM with probability \(T\) [2605.24914].

This deployment structure indicates that MVR-cache is intended as a drop-in semantic-caching layer for LLM APIs rather than a replacement for the cache controller itself. A plausible implication is that its compatibility with existing semantic-caching stacks depends chiefly on whether they can expose per-neighbor correctness metadata in the style of vCache.

## 6. Empirical results, comparisons, and scope

The experimental evaluation covers four datasets: SemCacheSearchQueries, SemCacheClassification, PromptBench (SQuAD-V2 perturbed), and QNLI-derived QA prompts [2605.24914]. Only about 3K prompts per dataset are used for training the segmentation model, which the paper presents as reflecting realistic labeling budgets [2605.24914]. Baselines include vCache, vCache with ColBERT-style segmentation, and vCache with POQD [2605.24914].

Across these datasets, the main reported result is that MVR-cache consistently increases cache hit rates by up to 37% while maintaining the same correctness guarantees [2605.24914]. More specifically, the paper reports cache-on-miss improvements up to about 25% and always-cache improvements up to 37%, notably on SemCacheSearchQueries [2605.24914]. Error rates for all methods converge below \(\delta\), and MVR-cache remains below the same bound while improving hit rate [2605.24914].

Latency results are also notable. Despite the additional segmentation and multi-vector reranking, end-to-end latency is reported to be reduced by up to 6% versus vCache at the same error bound, because the additional cache hits offset the retrieval overhead by reducing expensive LLM calls [2605.24914]. The paper characterizes the non-LLM overhead as tiny compared with LLM latency, with segmentation, embedding, and retrieval operating on the scale of tens of milliseconds versus thousands of milliseconds for an LLM call [2605.24914].

Ablations support several design choices. BGE, GTE, and E5 show negligible differences in performance within MVR-cache [2605.24914]. Training on 3K prompts performs nearly as well as 6K or 10K [2605.24914]. Candidate split positions based on punctuation perform similarly to token-level, keyword, and sentence-level candidates, and are described as slightly best [2605.24914]. Symmetric SMaxSim yields a measurable, though modest, improvement over unidirectional MaxSim [2605.24914].

The paper also reports cross-dataset transfer: a segmentation model trained on PromptBench transfers to QNLI and still outperforms vCache and the other baselines on hit rate without rebalance [2605.24914]. This suggests that the segmentation policy is not purely dataset-specific, although the paper still frames training as domain- and task-sensitive offline preparation.

The following table condenses the empirical claims emphasized in the paper.

| Aspect | Reported outcome | Source |
|---|---|---|
| Hit rate | Up to 37% increase | [2605.24914] |
| Correctness | Same guarantees as vCache; error below \(\delta\) | [2605.24914] |
| Latency | Up to 6% reduction vs vCache | [2605.24914] |

## 7. Interpretation, terminology, and limitations

In its primary usage, “MVR-cache” denotes the semantic caching system introduced in [2605.24914]. The acronym derives from Multi-Vector Retrieval, and the defining features are learned prompt segmentation, symmetric normalized MaxSim, and correctness-aware RL training. This is distinct from other uses of “MVR” in the literature, such as “Multi-View Video Reward Shaping” in reinforcement learning [2603.01694], and from separate multimodal KV-cache designs that have been described informally as MVR-style caches in vision or multimodal autoregressive settings [2505.19602, 2606.23581]. The shared acronym should not obscure the fact that [2605.24914] concerns semantic caching for LLM prompts rather than KV-cache reuse or video reward shaping.

Several limitations are explicit in the MVR-cache paper. The method requires labeled pairs or proxy labels for offline training; its formal guarantee relies on the Gaussian class-conditional similarity assumption; RL training is more complex than training-free vCache; and nearest-neighbor recomputation remains expensive even though it is amortized by periodic updates [2605.24914]. The current segmentation model is text-only, and the paper identifies multimodal prompts as future work [2605.24914].

These limitations clarify the scope of the contribution. MVR-cache does not replace correctness-aware semantic caching with a looser heuristic; rather, it attempts to improve the retrieval score while keeping the existing correctness framework intact. It also does not directly learn the embedding model jointly with the segmenter. The paper identifies direct multi-vector indexing, richer RL formulations such as actor-critic, and multimodal extensions as future directions [2605.24914].

Taken together, MVR-cache can be situated as a retrieval-centric redesign of semantic caching. Its novelty lies in treating cache hits as a structured retrieval problem over learned prompt segments, and in coupling that retrieval problem to a theorem-backed objective and an RL-trained segmentation policy.

Source: https://www.emergentmind.com/topics/mvr-cache