---
title: Visual Speech Deduplication in LLM Pipelines
url: https://www.emergentmind.com/topics/visual-speech-deduplication-strategy
type: topic
---

# Visual Speech Deduplication in LLM Pipelines

A visual speech deduplication strategy is a principled approach to reducing computational redundancy in visual speech processing pipelines by compacting feature representations based on clustered temporal similarity. In the context of the VSP-LLM framework, deduplication operates by collapsing contiguous runs of video frames mapped to identical “visual speech units”—phoneme-like discrete representations of the input latent space—thereby compressing the sequence length input to a large language model (LLM) without sacrificing recognition or translation accuracy. Empirical results demonstrate significant efficiency gains, with reductions in both the number of required floating-point operations (FLOPs) and memory consumption. This strategy integrates natively into pipelines leveraging Low-Rank Adaptation (LoRA), supporting scalable, context-aware visual speech recognition and translation [2402.15151].

## 1. Extraction and Definition of Visual Speech Units

The deduplication process begins with each raw video frame $x_t$ ($t = 1, \ldots, T$) processed by a pre-trained self-supervised visual encoder $f_{ss}$ (specifically, AV-HuBERT). This mapping produces a latent vector $z_t = f_{ss}(x_t) \in \mathbb{R}^d$. All latent representations $z_t$ in the training data are pooled and subjected to $K$-means clustering, yielding $K$ centroids $\{c_1,\dots, c_K\}$, each denoted a visual speech unit.

During both training and inference, each frame’s latent $z_t$ is assigned a unit index $u_t$ via
$$
u_t = \arg\min_{k \in \{1,\dots,K\}} \| z_t - c_k \|_2
$$
This procedure discretizes the sequence as $\{u_t\}_{t=1}^T$, where $u_t$ points to one of the $K$ visual speech units.

## 2. Redundancy Criteria and Deduplication Algorithm

Redundancy is characterized by the equivalence of adjacent unit assignments. No explicit similarity threshold is used; redundancy exists wherever $u_t = u_{t+1}$, so the binary indicator is
$$
s(t, t+1) = 
\begin{cases}
1 & \text{if } u_t = u_{t+1} \\
0 & \text{otherwise}
\end{cases}
$$
Contiguous runs of identical unit indices are grouped: for a run $i$ spanning frames $t_{i,\text{start}}$ to $t_{i,\text{end}}$, define $G_i = \{ t : t_{i,\text{start}} \leq t \leq t_{i,\text{end}}, u_t = \mathrm{constant} \}$. The deduplicated segment feature $\tilde{z}_i$ is the simple mean over corresponding latent vectors:
$$
\tilde{z}_i = \frac{1}{|G_i|} \sum_{t \in G_i} z_t
$$

The deduplicated sequence $\{\tilde{z}_i\}_{i=1}^{T'}$, where $T' \ll T$, replaces the original per-frame latent sequence in subsequent LLM input formation.

## 3. Workflow and Pseudocode

The deduplication pipeline is described in the following pseudocode:

```
Input: Raw video frames {x_t}_t=1..T
       Pre-trained visual encoder f_ss
       Cluster centroids {c_k}_k=1..K
Output: Deduplicated latent sequence {z̃_i}_i=1..T'

1. For t in 1..T:
       z_t = f_ss(x_t)              // d-dimensional
       u_t = argmin_k || z_t - c_k ||_2

2. Initialize segments = []
   curr_unit = u_1
   start = 1

3. For t from 2 to T:
       if u_t ≠ curr_unit:
          end = t - 1
          segments.append((start, end, curr_unit))
          start = t
          curr_unit = u_t
   // finish last segment
   segments.append((start, T, curr_unit))

4. For each segment i in segments:
       let G_i = {t : start_i ≤ t ≤ end_i}
       z̃_i = (1/|G_i|) * sum_{t∈G_i} z_t

5. Return {z̃_i}_i=1..T'
```

These averaged features are then mapped to the LLM token-embedding space via a linear transformation and concatenated with natural language instructions before being consumed by the LLM.

## 4. Computational Benefits and Empirical Results

Deduplication produces a compressed latent sequence, empirically reducing the average sequence length on the MuAViC benchmark by approximately $46.6\%$ when using $K=200$ clusters ($T' \approx 0.53 \cdot T$). This compression yields substantial computational savings:

| Setting                | No Deduplication | With Deduplication (K=200) |
|------------------------|------------------|----------------------------|
| FLOPs/train epoch      | 62.4 Peta        | 45.6 Peta (–26.9%)         |
| FLOPs/inference        | 19.2 Peta        | 14.0 Peta (–27.1%)         |
| Sequence length factor | 1.00             | 0.53                       |
| Avg BLEU               | 14.6             | 14.5                       |

No measurable loss in visual speech translation (BLEU: 14.6→14.5) or visual speech recognition (WER unchanged) was observed for $K=200$. Larger $K$ offer finer-grained units with less deduplication (higher fidelity, higher compute), while smaller $K$ lead to greater compression at the cost of marginal fidelity loss (e.g., $K=50$ gives up to $34\%$ FLOPs savings but minor BLEU drop to $14.3$) [2402.15151].

## 5. Integration with Large Language Models and LoRA

The deduplicated visual token sequence is linearly projected to match the LLM embedding space and combined with task instructions. The downstream LLM (LLaMA2-7B) is fine-tuned using QLoRA, updating only low-rank adapters. Since deduplication shortens the token sequence, each LoRA-adapted forward/backward pass processes fewer tokens, directly reducing per-step memory and compute requirements.

The number of clusters $K$ governs the trade-off between computational savings and output fidelity. In practice, $K=200$ provides an optimal trade-off for VSP-LLM, offering approximately 27% speedup with negligible BLEU or WER degradation.

## 6. Quantitative Benchmarks and Comparative Metrics

Experiments on MuAViC highlight the impact of deduplication within the VSP-LLM pipeline:

- Baseline (no deduplication): sequence length factor 1.00, FLOPs/train epoch 62.4 Peta, Avg BLEU 14.6.
- With deduplication ($K=200$): sequence length factor 0.53, FLOPs/train epoch 45.6 Peta, Avg BLEU 14.5.
- Full VSP-LLM (dedup + LoRA): for 30 hours labeled training, VST BLEU = 18.2 versus 19.2 for a cascaded AV-HuBERT+MT pipeline trained with 433 hours. VSR WER is 29.8% (30 h data) and 26.7% (433 h data).

A summary of the key quantitative results appears below:

| System                           | Training Data | VST BLEU | VSR WER  | FLOPs/train epoch |
|-----------------------------------|--------------|----------|----------|-------------------|
| Cascaded AV-HuBERT+MT            | 433 h        | 19.2     | —        | —                 |
| VSP-LLM (dedup+LoRA, K=200)      | 30 h         | 18.2     | 29.8%    | 45.6 Peta         |
| VSP-LLM (dedup+LoRA, K=200)      | 433 h        | —        | 26.7%    | —                 |

*This suggests that deduplication is highly effective at compressing visual speech feature streams for LLM-based recognition and translation, with minimal loss in performance and significant computational efficiency gains.*

## 7. Context, Significance, and Applicability

The deduplication strategy adopted in VSP-LLM introduces a lightweight, clustering-based compression mechanism tightly integrated with downstream LLM architectures and parameter-efficient adaptation (LoRA). Operating on the insight that contiguous frames with the same visual speech unit convey redundant information, deduplication exploits the temporal coherence of phoneme-like structures in visual speech data.

This approach is particularly significant for scalable, instruction-driven multi-task visual speech processing, making large multimodal models tractable for training and inference on long video sequences. It underscores the importance of temporal redundancy reduction when bridging raw modality streams with large, compute-hungry language models, without the need for hand-engineered similarity thresholds or segment boundaries. The method generalizes to visual speech recognition and translation, and is empirically validated as a robust, lossless compression technique in contemporary LLM pipelines [2402.15151].

Source: https://www.emergentmind.com/topics/visual-speech-deduplication-strategy