---
title: 'HyTIP: Hybrid Temporal Video Coding'
url: https://www.emergentmind.com/topics/hytip
type: topic
---

# HyTIP: Hybrid Temporal Video Coding

HyTIP, short for **Hybrid Temporal Information Propagation**, is a learned video coding framework for **masked conditional residual video coding** that combines two temporal propagation mechanisms usually treated separately in frame-based learned codecs: propagation of **explicit decoded signals** and propagation of **implicit latent features**. In the paper’s formulation, most frame-based learned video codecs can be interpreted as recurrent neural networks over time, and HyTIP is designed to improve the rate-distortion and buffer-size trade-off within that RNN view. It uses a **hybrid buffering strategy** in both motion coding and inter-frame coding, achieving competitive coding performance with a much smaller buffer size; the reported comparisons state that it outperforms **VTM 17.0 (Low-delay B)** in **PSNR-RGB** and **MS-SSIM-RGB** [2508.02072].

## 1. RNN formulation and the problem HyTIP addresses

HyTIP is motivated by a reinterpretation of learned video coding as temporal recurrence. Under this view, two dominant propagation strategies are identified. The first is **output-recurrence**, also described as **explicit buffering**, in which the codec propagates the previously decoded frame \(\hat{x}_{t-1}\). This is intuitive and closely aligned with classical predictive video coding, but it imposes a **dual constraint** on the decoded frame: it must both approximate the true frame well and carry useful temporal information for future coding. The second is **hidden-to-hidden recurrence**, also described as **implicit buffering**, in which the codec propagates latent feature maps \(F_{t-1}\). This is more flexible because the buffered representation is not constrained to resemble a frame, but it tends to require large, non-compact, memory-heavy buffers [2508.02072].

HyTIP is introduced as a response to these limitations. Its central design claim is that a learned codec need not choose between decoded-frame recurrence and latent-feature recurrence. Instead, it stores both: an explicit, semantically meaningful decoded reference and a small number of implicit latent features containing complementary temporal information. In the paper’s RNN terminology, this combines **output-to-output recurrence** and **hidden-to-hidden recurrence** in a single codec.

## 2. Architecture and per-frame coding pipeline

For each time step \(t\), HyTIP follows a sequential pipeline. First, it performs **motion estimation** between the current frame \(x_t\) and the reference frame \(\hat{x}_{t-1}\) using

\[
f_t = \text{MENet}(x_t,\hat{x}_{t-1}).
\]

Second, it performs **motion coding**, encoding and transmitting the flow \(f_t\) to obtain decoded flow \(\hat{f}_t\). Third, it performs **temporal prediction** by motion-compensating buffered references with \(\hat{f}_t\). These references include both the explicit decoded frame \(\hat{x}_{t-1}\) and the implicit latent features \(F_{t-1}\). This stage produces a pixel-domain predictor \(x_c\) and multi-scale feature predictors \(\{C_1,C_2,C_3\}\). Fourth, it performs **inter-frame coding**, where the residual is coded in a masked conditional residual coding framework using contextual encoder and decoder modules denoted \(G^{enc}\) and \(G^{dec}\), conditioned on \(\{C_1,C_2,C_3\}\). Fifth, it performs a **buffer update**, storing the new decoded frame and selected latent features for the next time step [2508.02072].

The same hybrid principle is extended to motion coding. HyTIP buffers the decoded previous flow \(\hat{f}_{t-1}\) as an explicit motion reference and motion latent features \(F^f_{t-1}\) as an implicit motion reference. Unlike inter-frame references, these motion references are used **without motion compensation**. The motion latent features are stored at **one-fourth resolution** in width and height.

## 3. Hybrid buffering as the defining mechanism

The defining feature of HyTIP is its **hybrid buffering strategy**. In inter-frame coding, the buffer contains the previously decoded frame \(\hat{x}_{t-1}\) and latent features \(F_{t-1}\). In motion coding, it contains decoded previous flow \(\hat{f}_{t-1}\) and motion latent features \(F^f_{t-1}\). The paper emphasizes that the decoded frame is the **primary source** of temporal information, while latent features are **complementary**. This division of labor is intended to reduce the burden placed on latent buffers while preserving the flexibility absent from explicit-only schemes [2508.02072].

Within the paper’s interpretation, output recurrence is limited because the decoded frame is forced to satisfy its reconstruction role and its temporal-propagation role simultaneously. Hidden recurrence avoids that dual-role constraint, but large feature buffers become a practical liability. HyTIP’s hybrid design is presented as a compromise that preserves compactness and interpretability while retaining part of the expressive advantage of hidden recurrence. The paper characterizes this as the **first masked conditional residual coding with hybrid buffering**.

## 4. Objective functions and training protocol

HyTIP is trained with rate-distortion objectives, but the loss is decomposed by stage. For motion coding with explicit reference, the paper gives

\[
R^{motion}_t + \lambda \times D(x_t, warp(x_{t-1}, \hat{f}_t)).
\]

For motion compensation and predictor training, it uses

\[
\lambda \times D(x_t, x_c),
\]

and, in a later hybrid refinement stage,

\[
R_t + \lambda \times \frac{D(x_t, x_c) + D(x_t, \hat{x}_t)}{2}.
\]

For inter-frame coding, the objective is

\[
R_t + \lambda \times D(x_t, \hat{x}_t).
\]

Across training phases, the overall objective remains the standard rate-distortion form

\[
R_t + \lambda D(x_t,\hat{x}_t).
\]

The training schedule includes motion coding, context modules, inter-frame codec training, and later **variable-rate and EPA-like refinements** [2508.02072].

The reported data regimen is split between **pretraining** and **fine-tuning**. Pretraining uses **Vimeo-90K**, specified as **91,701 7-frame sequences**. Fine-tuning uses **BVI-DVC**, specified as **800 64-frame sequences**. The training regime begins with **5-frame training** and then uses **10-frame training** for longer-sequence fine-tuning.

## 5. Evaluation protocol and ablation findings

Evaluation is reported on **UVG**, **MCL-JCV**, **HEVC Classes B–E**, and **HEVC-RGB**. The protocol encodes the first **96 frames** with **intra period = 32**. Input frames are converted from **YUV420 to RGB444 using BT.601**, and supplementary comparisons are also reported under **BT.709**. The reported metrics are **PSNR-RGB**, **MS-SSIM-RGB**, bitrate in **bpp**, and **BD-rate**, where positive BD-rate means bitrate inflation and negative BD-rate means bitrate savings [2508.02072].

The paper’s central ablation compares **explicit buffering**, **implicit buffering**, and **hybrid buffering**. The reported pattern is that implicit buffering with a large buffer outperforms explicit buffering, but degrades sharply when the buffer size is reduced; hybrid buffering remains robust even with small buffers. Representative results from the reported table use explicit buffering as the anchor. For **motion coding + explicit inter-frame coding**, the paper reports **-12.2% average BD-rate** for implicit motion with large buffer, **-16.0%** for hybrid motion with large buffer, and **-14.7%** for hybrid motion with small buffer. For hybrid settings with implicit inter-frame coding and small buffer, the reported values are **-15.0%** for **Hybrid (motion 2+0.125, inter 0+5)** and **-21.5%** for **Hybrid (motion 2+0.125, inter 2+3)**.

Longer-sequence training improves all methods, but not equally. Under **5-frame vs 10-frame training**, the paper reports improvement of about **3.2%** for explicit buffering, **5.8%** for implicit buffering, and **5.2%** for hybrid buffering. The paper interprets this as evidence that explicit recurrence remains constrained by the decoded-frame bottleneck, whereas hidden and hybrid schemes exploit longer temporal context more effectively.

## 6. Comparative performance, buffer size, and reported significance

A major practical claim of HyTIP is that it improves the rate-distortion and memory trade-off. The paper reports the following buffer sizes for learned P-frame codecs:

| Codec | Buffer size |
|---|---:|
| DCVC-TCM | 67 |
| DCVC-HEM | 67.625 |
| DCVC-DC | 55.75 |
| DCVC-FM | 55.75 |
| MaskCRT | 13 |
| HyTIP | 7.875 |

This makes HyTIP the **smallest buffer size** among the compared learned P-frame codecs. The paper further states that HyTIP uses only about **14% of the buffer size** required by implicit-buffer methods such as **DCVC-TCM**, **DCVC-DC**, and **DCVC-FM** [2508.02072].

Against **VTM 17.0**, the average **BT.601 PSNR-RGB BD-rate** values reported are **+31.5%** for **HM 16.25**, **+4.1%** for **MaskCRT**, **+33.9%** for **DCVC-TCM**, **-5.8%** for **DCVC-HEM**, **-23.6%** for **DCVC-DC**, **-18.5%** for **DCVC-FM**, and **-22.1%** for **HyTIP**. For **BT.601 MS-SSIM-RGB**, the reported averages are **+27.2%** for **HM**, **-39.0%** for **MaskCRT**, **-26.4%** for **DCVC-TCM**, **-47.0%** for **DCVC-HEM**, **-55.2%** for **DCVC-DC**, and **-56.7%** for **HyTIP**. Under **BT.709**, the reported average **PSNR-RGB BD-rate** is **-17.4%** for HyTIP, while the **MS-SSIM-RGB BD-rate** is **-53.8%**. The paper notes that training procedures differ across methods and that some public test models were used, so direct one-to-one comparison requires caution.

The paper’s stated contributions are fourfold: an **RNN perspective on learned video coding**; the **first masked conditional residual coding with hybrid buffering**; **strong efficiency**, including comparable performance to state-of-the-art learned codecs with much smaller buffer size; and a **generalizable design** that can be extended to other learned video codecs. The stated future direction is to improve functionality, including **YUV coding** and broader practical coding settings. The paper also notes that its main focus is **temporal propagation design**, not a full exploration of every codec component, and that additional optimization could potentially reduce model size or compute further. The source code is reported as available at **https://github.com/NYCU-MAPL/HyTIP**.

Source: https://www.emergentmind.com/topics/hytip