Papers
Topics
Authors
Recent
Search
2000 character limit reached

Minimal Predictive Sufficiency SSM

Updated 26 February 2026
  • The paper introduces an info-theoretic framework where the hidden state is a minimal predictive sufficient statistic for accurate future forecasting.
  • It employs a relaxed Lagrangian objective combining prediction loss with an information regularizer to compress non-causal history efficiently.
  • Empirical results show MPS-SSM achieves state-of-the-art accuracy and robustness against noisy inputs across multiple time-series benchmarks.

The Minimal Predictive Sufficiency State Space Model (MPS-SSM) is a sequence modeling framework whose content-selective state gating is derived from a first-principle information-theoretic criterion. MPS-SSM builds on the principle that the model’s hidden state should be a minimal sufficient statistic of the past for predicting the future. This results in a model that maximally compresses historical context, learns to ignore non-causal information, and exhibits robustness and accuracy across long-horizon and noisy forecasting scenarios (Wang et al., 5 Aug 2025).

1. Principle of Predictive Sufficiency

The central theoretical construct underlying MPS-SSM is the principle of predictive sufficiency. For a sequence (U1:t,Yt:t+τ)(U_{1:t}, Y_{t:t+\tau}) where U1:tU_{1:t} is the observed history and Yt:t+τY_{t:t+\tau} denotes a segment of future targets, MPS-SSM demands that the hidden state hth_t at every time tt satisfies two criteria:

  • Predictive sufficiency: The hidden state hth_t must retain all information in U1:tU_{1:t} relevant for predicting Yt:t+τY_{t:t+\tau}; formally, I(ht;Yt:t+τ)=I(U1:t;Yt:t+τ)I(h_t;Y_{t:t+\tau}) = I(U_{1:t};Y_{t:t+\tau}).
  • Minimality: Among all statistics satisfying sufficiency, hth_t should minimize U1:tU_{1:t}0, i.e., U1:tU_{1:t}1 for any U1:tU_{1:t}2 also satisfying sufficiency.

Collectively, these constraints characterize U1:tU_{1:t}3 as a minimal predictive sufficient statistic and can be formalized by the optimization problem: U1:tU_{1:t}4 This setup ensures that the hidden state captures only the causal structure necessary for accurate sequence prediction and discards spurious or non-predictive variability.

2. MPS-SSM Objective Function Derivation

Directly enforcing the constraint U1:tU_{1:t}5 is intractable, so MPS-SSM introduces a relaxed Lagrangian objective. The predictive sufficiency criterion is represented by a standard prediction loss,

U1:tU_{1:t}6

while the minimality term is realized as an information-theoretic regularizer,

U1:tU_{1:t}7

The total objective becomes

U1:tU_{1:t}8

with U1:tU_{1:t}9 balancing prediction performance and information compression.

As direct computation of Yt:t+τY_{t:t+\tau}0 is intractable, MPS-SSM employs a variational upper bound using a decoder Yt:t+τY_{t:t+\tau}1: Yt:t+τY_{t:t+\tau}2 enabling practical and stable optimization via backpropagation with

Yt:t+τY_{t:t+\tau}3

3. Architecture and Training Methodology

MPS-SSM extends a content-selective SSM backbone—such as Mamba—by integrating several key modules:

  • Selection Gate Yt:t+τY_{t:t+\tau}4: Computes adaptive state-space parameters Yt:t+τY_{t:t+\tau}5 conditioned on each input Yt:t+τY_{t:t+\tau}6.
  • SSM Recurrence: The core transition follows

Yt:t+τY_{t:t+\tau}7

where Yt:t+τY_{t:t+\tau}8 is approximated via zero-order hold or NPLR techniques.

  • Minimality Module: A lightweight decoder Yt:t+τY_{t:t+\tau}9 reconstructs hth_t0 from hth_t1 to facilitate the variational information regularization.
  • Prediction Head: Projects hth_t2 into target predictions hth_t3.

Training is conducted over entire unrolled sequences, jointly optimizing hth_t4 (gate), hth_t5 (decoder), and SSM matrices to minimize hth_t6 with standard first-order methods. The entire process is efficiently scalable and practical for large-scale time-series tasks.

Training Workflow Table

Step Operation Output
Selection Gate hth_t7 Adaptive params
SSM Recurrence hth_t8 Hidden state
Prediction hth_t9 Forecasted values
Minimality Module tt0 Info loss
Backpropagation tt1 Parameter update

4. Empirical Results and Robustness Analysis

MPS-SSM has been evaluated on established sequence modeling and forecasting benchmarks, including ETT (ETTh1/2, ETTm1/2), Weather, Electricity, Traffic, and Exchange, across forecast horizons (96, 192, 336, 720) and measured via MSE and MAE.

Key findings include:

  • Optimal Regularization (tt2) Sensitivity: Each dataset and horizon displays a “sweet-spot” tt3 (e.g., ETTh1: tt4; Weather: tt5; ETTm2: tt6), and the optimal tt7 increases with forecast length.
  • State-of-the-Art Accuracy:
    • On ETTh1 (96), MPS-SSM achieves MSE = 0.375, second only to PatchTST (0.360).
    • On ETTm2 (96), MSE = 0.165, outperforming PatchTST (0.224).
    • On Electricity (96), MSE = 0.151 (vs. next-best 0.225).
    • On long horizons, MPS-SSM routinely ranks best or second-best.
  • Robustness to Noise: Under impulse noise perturbations to inputs, increasing tt8 monotonically reduces forecast error degradation; at tt9, degradation is approximately threefold lower than at hth_t0. This empirically validates the theoretical prediction that MPS-SSM is resilient to non-causal spurious input patterns.

5. Generalization to a Regularization Framework

The MPS principle is not restricted to SSMs and can be instantiated as a model-agnostic regularizer for any sequential architecture. This extension involves:

  1. Selecting an internal representation hth_t1 (e.g., an SSM state, Transformer embedding, or linear hidden vector).
  2. Attaching a lightweight decoder hth_t2.
  3. Adding the minimality regularization term hth_t3 to the base task loss.

This general regularization strategy takes the form: hth_t4 where hth_t5 denotes the original task model.

Empirical evidence demonstrates utility across architectures such as Mamba (MPS-Mamba), linear models (MPS-DLinear), and Transformers (MPS-PatchTST), with consistent improvements on ETT and other datasets (e.g., MPS-PatchTST achieves ETTh1/96 MSE=0.328 compared to 0.360 for vanilla PatchTST).

6. Significance and Implications

MPS-SSM is the first selective SSM whose gating is derived from the information-theoretic requirement that hidden states encode the minimal predictive sufficient statistic. The resulting mutual information penalty confers both empirical state-of-the-art generalization and robustness properties, notably resistance to non-causal and spurious noise. Furthermore, the principle’s generality enables its adoption as an effective regularizer in architectures beyond SSMs, including popular sequence models such as Transformers and linear models (Wang et al., 5 Aug 2025). A plausible implication is the emergence of a new paradigm for designing sequential models grounded in first principles rather than heuristic mechanism design.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Minimal Predictive Sufficiency SSM (MPS-SSM).