Papers
Topics
Authors
Recent
Search
2000 character limit reached

HyTIP: Hybrid Temporal Video Coding

Updated 7 July 2026
  • HyTIP is a learned video coding framework that integrates explicit decoded frame propagation with implicit latent feature buffering to enhance temporal coding.
  • It employs a hybrid buffering strategy in both motion and inter-frame coding, optimizing the rate-distortion and memory trade-offs with a significantly reduced buffer size.
  • The framework outperforms state-of-the-art codecs like VTM 17.0 in PSNR-RGB and MS-SSIM-RGB metrics while using only about 14% of the typical buffer size required by implicit methods.

HyTIP, short for Hybrid Temporal Information Propagation, is a learned video coding framework for masked conditional residual video coding that combines two temporal propagation mechanisms usually treated separately in frame-based learned codecs: propagation of explicit decoded signals and propagation of implicit latent features. In the paper’s formulation, most frame-based learned video codecs can be interpreted as recurrent neural networks over time, and HyTIP is designed to improve the rate-distortion and buffer-size trade-off within that RNN view. It uses a hybrid buffering strategy in both motion coding and inter-frame coding, achieving competitive coding performance with a much smaller buffer size; the reported comparisons state that it outperforms VTM 17.0 (Low-delay B) in PSNR-RGB and MS-SSIM-RGB (Chen et al., 4 Aug 2025).

1. RNN formulation and the problem HyTIP addresses

HyTIP is motivated by a reinterpretation of learned video coding as temporal recurrence. Under this view, two dominant propagation strategies are identified. The first is output-recurrence, also described as explicit buffering, in which the codec propagates the previously decoded frame x^t−1\hat{x}_{t-1}. This is intuitive and closely aligned with classical predictive video coding, but it imposes a dual constraint on the decoded frame: it must both approximate the true frame well and carry useful temporal information for future coding. The second is hidden-to-hidden recurrence, also described as implicit buffering, in which the codec propagates latent feature maps Ft−1F_{t-1}. This is more flexible because the buffered representation is not constrained to resemble a frame, but it tends to require large, non-compact, memory-heavy buffers (Chen et al., 4 Aug 2025).

HyTIP is introduced as a response to these limitations. Its central design claim is that a learned codec need not choose between decoded-frame recurrence and latent-feature recurrence. Instead, it stores both: an explicit, semantically meaningful decoded reference and a small number of implicit latent features containing complementary temporal information. In the paper’s RNN terminology, this combines output-to-output recurrence and hidden-to-hidden recurrence in a single codec.

2. Architecture and per-frame coding pipeline

For each time step tt, HyTIP follows a sequential pipeline. First, it performs motion estimation between the current frame xtx_t and the reference frame x^t−1\hat{x}_{t-1} using

ft=MENet(xt,x^t−1).f_t = \text{MENet}(x_t,\hat{x}_{t-1}).

Second, it performs motion coding, encoding and transmitting the flow ftf_t to obtain decoded flow f^t\hat{f}_t. Third, it performs temporal prediction by motion-compensating buffered references with f^t\hat{f}_t. These references include both the explicit decoded frame x^t−1\hat{x}_{t-1} and the implicit latent features Ft−1F_{t-1}0. This stage produces a pixel-domain predictor Ft−1F_{t-1}1 and multi-scale feature predictors Ft−1F_{t-1}2. Fourth, it performs inter-frame coding, where the residual is coded in a masked conditional residual coding framework using contextual encoder and decoder modules denoted Ft−1F_{t-1}3 and Ft−1F_{t-1}4, conditioned on Ft−1F_{t-1}5. Fifth, it performs a buffer update, storing the new decoded frame and selected latent features for the next time step (Chen et al., 4 Aug 2025).

The same hybrid principle is extended to motion coding. HyTIP buffers the decoded previous flow Ft−1F_{t-1}6 as an explicit motion reference and motion latent features Ft−1F_{t-1}7 as an implicit motion reference. Unlike inter-frame references, these motion references are used without motion compensation. The motion latent features are stored at one-fourth resolution in width and height.

3. Hybrid buffering as the defining mechanism

The defining feature of HyTIP is its hybrid buffering strategy. In inter-frame coding, the buffer contains the previously decoded frame Ft−1F_{t-1}8 and latent features Ft−1F_{t-1}9. In motion coding, it contains decoded previous flow tt0 and motion latent features tt1. The paper emphasizes that the decoded frame is the primary source of temporal information, while latent features are complementary. This division of labor is intended to reduce the burden placed on latent buffers while preserving the flexibility absent from explicit-only schemes (Chen et al., 4 Aug 2025).

Within the paper’s interpretation, output recurrence is limited because the decoded frame is forced to satisfy its reconstruction role and its temporal-propagation role simultaneously. Hidden recurrence avoids that dual-role constraint, but large feature buffers become a practical liability. HyTIP’s hybrid design is presented as a compromise that preserves compactness and interpretability while retaining part of the expressive advantage of hidden recurrence. The paper characterizes this as the first masked conditional residual coding with hybrid buffering.

4. Objective functions and training protocol

HyTIP is trained with rate-distortion objectives, but the loss is decomposed by stage. For motion coding with explicit reference, the paper gives

tt2

For motion compensation and predictor training, it uses

tt3

and, in a later hybrid refinement stage,

tt4

For inter-frame coding, the objective is

tt5

Across training phases, the overall objective remains the standard rate-distortion form

tt6

The training schedule includes motion coding, context modules, inter-frame codec training, and later variable-rate and EPA-like refinements (Chen et al., 4 Aug 2025).

The reported data regimen is split between pretraining and fine-tuning. Pretraining uses Vimeo-90K, specified as 91,701 7-frame sequences. Fine-tuning uses BVI-DVC, specified as 800 64-frame sequences. The training regime begins with 5-frame training and then uses 10-frame training for longer-sequence fine-tuning.

5. Evaluation protocol and ablation findings

Evaluation is reported on UVG, MCL-JCV, HEVC Classes B–E, and HEVC-RGB. The protocol encodes the first 96 frames with intra period = 32. Input frames are converted from YUV420 to RGB444 using BT.601, and supplementary comparisons are also reported under BT.709. The reported metrics are PSNR-RGB, MS-SSIM-RGB, bitrate in bpp, and BD-rate, where positive BD-rate means bitrate inflation and negative BD-rate means bitrate savings (Chen et al., 4 Aug 2025).

The paper’s central ablation compares explicit buffering, implicit buffering, and hybrid buffering. The reported pattern is that implicit buffering with a large buffer outperforms explicit buffering, but degrades sharply when the buffer size is reduced; hybrid buffering remains robust even with small buffers. Representative results from the reported table use explicit buffering as the anchor. For motion coding + explicit inter-frame coding, the paper reports -12.2% average BD-rate for implicit motion with large buffer, -16.0% for hybrid motion with large buffer, and -14.7% for hybrid motion with small buffer. For hybrid settings with implicit inter-frame coding and small buffer, the reported values are -15.0% for Hybrid (motion 2+0.125, inter 0+5) and -21.5% for Hybrid (motion 2+0.125, inter 2+3).

Longer-sequence training improves all methods, but not equally. Under 5-frame vs 10-frame training, the paper reports improvement of about 3.2% for explicit buffering, 5.8% for implicit buffering, and 5.2% for hybrid buffering. The paper interprets this as evidence that explicit recurrence remains constrained by the decoded-frame bottleneck, whereas hidden and hybrid schemes exploit longer temporal context more effectively.

6. Comparative performance, buffer size, and reported significance

A major practical claim of HyTIP is that it improves the rate-distortion and memory trade-off. The paper reports the following buffer sizes for learned P-frame codecs:

Codec Buffer size
DCVC-TCM 67
DCVC-HEM 67.625
DCVC-DC 55.75
DCVC-FM 55.75
MaskCRT 13
HyTIP 7.875

This makes HyTIP the smallest buffer size among the compared learned P-frame codecs. The paper further states that HyTIP uses only about 14% of the buffer size required by implicit-buffer methods such as DCVC-TCM, DCVC-DC, and DCVC-FM (Chen et al., 4 Aug 2025).

Against VTM 17.0, the average BT.601 PSNR-RGB BD-rate values reported are +31.5% for HM 16.25, +4.1% for MaskCRT, +33.9% for DCVC-TCM, -5.8% for DCVC-HEM, -23.6% for DCVC-DC, -18.5% for DCVC-FM, and -22.1% for HyTIP. For BT.601 MS-SSIM-RGB, the reported averages are +27.2% for HM, -39.0% for MaskCRT, -26.4% for DCVC-TCM, -47.0% for DCVC-HEM, -55.2% for DCVC-DC, and -56.7% for HyTIP. Under BT.709, the reported average PSNR-RGB BD-rate is -17.4% for HyTIP, while the MS-SSIM-RGB BD-rate is -53.8%. The paper notes that training procedures differ across methods and that some public test models were used, so direct one-to-one comparison requires caution.

The paper’s stated contributions are fourfold: an RNN perspective on learned video coding; the first masked conditional residual coding with hybrid buffering; strong efficiency, including comparable performance to state-of-the-art learned codecs with much smaller buffer size; and a generalizable design that can be extended to other learned video codecs. The stated future direction is to improve functionality, including YUV coding and broader practical coding settings. The paper also notes that its main focus is temporal propagation design, not a full exploration of every codec component, and that additional optimization could potentially reduce model size or compute further. The source code is reported as available at https://github.com/NYCU-MAPL/HyTIP.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HyTIP.