Papers
Topics
Authors
Recent
Search
2000 character limit reached

Q-MMR: Moment-Matched Reweighting

Updated 4 July 2026
  • Q-MMR is a reweighting framework that applies moment-matched constraints to estimate expected returns from off-policy data using learned scalar weights.
  • It employs an inductive top‐down moment matching against value-function discriminators, shifting estimation from density ratios to local compatibility constraints.
  • The finite-sample guarantee is dimension‐free and data‐dependent, recovering established estimators like Monte Carlo, step-wise importance sampling, and linear FQE.

Q-MMR, or Q-function based Moment-Matched Reweighting, is a theoretical framework for off-policy evaluation in finite-horizon Markov decision processes. Introduced by Xiang Li and Nan Jiang, it learns one scalar weight for each data point so that reweighted rewards approximate the expected return of a target policy. Its defining mechanism is an inductive, top-down moment-matching procedure against a value-function discriminator class, and its main theoretical result is a data-dependent finite-sample guarantee under only the realizability of QπQ^\pi, with a dimension-free leading term that does not depend on the statistical complexity of the function class (Li et al., 7 May 2026).

1. Formal setting and estimation target

Q-MMR is formulated for a finite-horizon MDP with disjoint state layers S0,…,SHS_0,\ldots,S_H, action set AA, transition kernel P(s′∣s,a)P(s' \mid s,a) supported on the next layer, and reward distribution R(r∣s,a)R(r \mid s,a) with r∈[0,Rmax⁡]r \in [0,R_{\max}]. The setup includes a fixed initial dummy pair (s0,a0)(s_0,a_0) and reward r0≡0r_0 \equiv 0. For a target policy π:S→Δ(A)\pi : S \to \Delta(A), trajectories have the form

(s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),

with occupancy

S0,…,SHS_0,\ldots,S_H0

and return

S0,…,SHS_0,\ldots,S_H1

The available data are S0,…,SHS_0,\ldots,S_H2 i.i.d. trajectories collected under a behavior policy S0,…,SHS_0,\ldots,S_H3, written as

S0,…,SHS_0,\ldots,S_H4

with empirical marginal S0,…,SHS_0,\ldots,S_H5 and expectation S0,…,SHS_0,\ldots,S_H6. The off-policy evaluation objective is to estimate S0,…,SHS_0,\ldots,S_H7 using the behavior data and a function class S0,…,SHS_0,\ldots,S_H8 under the standard realizability assumption that, for each S0,…,SHS_0,\ldots,S_H9, the true action-value function

AA0

lies in AA1 (Li et al., 7 May 2026).

The estimator used by Q-MMR has the form

AA2

where the AA3 are learned scalar weights. The intuitive comparison point is the marginalized importance ratio AA4, which would yield unbiased reweighting if known. Q-MMR replaces explicit density-ratio estimation by a recursive empirical moment-matching construction (Li et al., 7 May 2026).

2. Recursive moment-matched reweighting

For each level AA5, given the previous weights AA6, Q-MMR defines the empirical matching loss

AA7

where

AA8

Q-MMR chooses AA9 to approximately minimize P(s′∣s,a)P(s' \mid s,a)0, and among minimizers selects the one with smallest P(s′∣s,a)P(s' \mid s,a)1-norm P(s′∣s,a)P(s' \mid s,a)2. Because P(s′∣s,a)P(s' \mid s,a)3 is convex, for example in the linear case, this is a convex-concave minimax problem in P(s′∣s,a)P(s' \mid s,a)4 (Li et al., 7 May 2026).

The recursion is initialized with P(s′∣s,a)P(s' \mid s,a)5 for all P(s′∣s,a)P(s' \mid s,a)6. For P(s′∣s,a)P(s' \mid s,a)7, the algorithm solves

P(s′∣s,a)P(s' \mid s,a)8

optionally subject to P(s′∣s,a)P(s' \mid s,a)9, and then forms the final estimate R(r∣s,a)R(r \mid s,a)0 by summing the reweighted rewards across levels. Although the procedure never explicitly computes R(r∣s,a)R(r \mid s,a)1, the analysis shows that each matching step controls the Bellman-error term

R(r∣s,a)R(r \mid s,a)2

through the requirement that, for all R(r∣s,a)R(r \mid s,a)3,

R(r∣s,a)R(r \mid s,a)4

This induces a top-down propagation of Bellman corrections from one level to the next (Li et al., 7 May 2026).

A plausible implication is that Q-MMR shifts the main estimation burden from explicit transition or occupancy modeling to a sequence of local compatibility constraints with the discriminator classes R(r∣s,a)R(r \mid s,a)5. In the terminology of the paper, the method learns weights inductively in a top-down manner via moment matching against a value-function discriminator class (Li et al., 7 May 2026).

3. Finite-sample theory and dimension-free control

The central finite-sample result states that, with probability at least R(r∣s,a)R(r \mid s,a)6, for any choice of weights R(r∣s,a)R(r \mid s,a)7 such that each depends only on the first R(r∣s,a)R(r \mid s,a)8 levels of data,

R(r∣s,a)R(r \mid s,a)9

where

r∈[0,Rmax⁡]r \in [0,R_{\max}]0

The leading term depends on the empirical matching losses r∈[0,Rmax⁡]r \in [0,R_{\max}]1, not on the covering number, dimension, or other standard complexity measure of r∈[0,Rmax⁡]r \in [0,R_{\max}]2. This is the sense in which the guarantee is dimension-free (Li et al., 7 May 2026).

The proof sketch in the paper proceeds by a Bellman-residual telescoping argument. First, r∈[0,Rmax⁡]r \in [0,R_{\max}]3 is replaced by the Bellman difference r∈[0,Rmax⁡]r \in [0,R_{\max}]4, with the resulting empirical-variance error bounded by r∈[0,Rmax⁡]r \in [0,R_{\max}]5. Second, the matching loss r∈[0,Rmax⁡]r \in [0,R_{\max}]6 is used to replace r∈[0,Rmax⁡]r \in [0,R_{\max}]7 by r∈[0,Rmax⁡]r \in [0,R_{\max}]8. Third, the resulting expression telescopes exactly to r∈[0,Rmax⁡]r \in [0,R_{\max}]9. The final error is therefore the sum of the moment-matching losses and the higher-order (s0,a0)(s_0,a_0)0-norm penalties (Li et al., 7 May 2026).

The paper emphasizes three consequences of this result. First, the guarantee requires only realizability of (s0,a0)(s_0,a_0)1, not full Bellman completeness. Second, the bound is data-dependent: the quantities (s0,a0)(s_0,a_0)2 can be computed on hold-out data, yielding what the authors describe as a “WYSIWYG” uncertainty measure. Third, the complexity of the function class enters only through the higher-order norm terms rather than the leading approximation term (Li et al., 7 May 2026).

4. Relation to importance sampling, FQE, and tabular OPE

Q-MMR is explicitly positioned as a unifying reweighting framework with connections to several established off-policy evaluation methods. In the on-policy case, when (s0,a0)(s_0,a_0)3 and one chooses (s0,a0)(s_0,a_0)4, the matching losses satisfy (s0,a0)(s_0,a_0)5, and (s0,a0)(s_0,a_0)6 reduces to the Monte Carlo estimator with a dimension-free Hoeffding bound. For step-wise importance sampling, if the second term of (s0,a0)(s_0,a_0)7 is modified to use (s0,a0)(s_0,a_0)8, then the choice

(s0,a0)(s_0,a_0)9

makes r0≡0r_0 \equiv 00, but with exponentially large variance. In this comparison, Q-MMR trades variance against exact matching by constraining the moments only through the discriminator classes r0≡0r_0 \equiv 01 (Li et al., 7 May 2026).

When r0≡0r_0 \equiv 02 is linear,

r0≡0r_0 \equiv 03

the paper states that Q-MMR with exact matching and minimum-r0≡0r_0 \equiv 04 weights exactly reproduces linear FQE. In the tabular case, Q-MMR coincides with computing empirical occupancy-ratio weights

r0≡0r_0 \equiv 05

and therefore with certainty-equivalence model-based evaluation (Li et al., 7 May 2026).

Method Q-MMR specialization Consequence
On-policy Monte Carlo r0≡0r_0 \equiv 06, r0≡0r_0 \equiv 07 r0≡0r_0 \equiv 08, recovers MC
Step-wise Importance Sampling Product ratios r0≡0r_0 \equiv 09 π:S→Δ(A)\pi : S \to \Delta(A)0, exponentially large variance
Linear FQE Linear π:S→Δ(A)\pi : S \to \Delta(A)1, exact matching, min-π:S→Δ(A)\pi : S \to \Delta(A)2 weights Exactly reproduces linear FQE
Tabular OPE Empirical occupancy-ratio weights Coincides with certainty-equivalence model-based evaluation

This structure suggests that Q-MMR is best understood not as a competitor to a single baseline, but as a reweighting template whose limiting cases recover several familiar estimators. The paper also notes a conceptual link to joint multi-level reweighting methods such as DualDICE and MWL, but frames simultaneous multi-level minimax optimization as difficult in the non-parametric regime because of double-sampling issues (Li et al., 7 May 2026).

5. Coverage, population weights, and projected feature geometry

A substantial part of the theory concerns the nature of coverage. The population weight functions π:S→Δ(A)\pi : S \to \Delta(A)3 are defined as minimizers of the population matching loss

π:S→Δ(A)\pi : S \to \Delta(A)4

When π:S→Δ(A)\pi : S \to \Delta(A)5 is a reproducing-kernel or linear class, the paper gives the representation

π:S→Δ(A)\pi : S \to \Delta(A)6

with

π:S→Δ(A)\pi : S \to \Delta(A)7

The associated coverage parameter is

π:S→Δ(A)\pi : S \to \Delta(A)8

Under Bellman completeness, π:S→Δ(A)\pi : S \to \Delta(A)9 (Li et al., 7 May 2026).

The interpretation given in the paper is that (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),0 measures how well the dataset covers the “projected next-step features” needed to evaluate the target policy. This reformulates coverage away from raw occupancy overlap and toward a feature-space quantity determined by the discriminator classes. The same parameter controls higher-order terms in the random-design bound and the sample size required for (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),1 with high probability (Li et al., 7 May 2026).

A plausible implication is that Q-MMR distinguishes between two kinds of mismatch: lack of support in the raw state-action space, and lack of support only after projection through (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),2. The framework treats the latter as the operative obstruction for estimation error, which is consistent with its reliance on moment constraints rather than full density-ratio recovery.

6. Advantages, limitations, and open problems

The advantages emphasized for Q-MMR are tightly tied to its theorem. The leading term in the finite-sample bound is dimension-free; the uncertainty decomposition is data-dependent and can be monitored on hold-out data; the framework recovers Monte Carlo, importance sampling, linear FQE, and tabular certainty-equivalence as special cases; and the main guarantee requires only realizability of (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),3, not Bellman completeness (Li et al., 7 May 2026).

The limitations are equally explicit. The current analysis assumes i.i.d. trajectories; adaptive or online data collection violates the conditional-independence structure used in the concentration arguments. Extending the framework to Markov-chain samplers or infinite-horizon discounted MDPs is identified as open. The algorithm is also greedy and local in the sense that each (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),4 optimizes only the level-(s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),5 matching objective without directly accounting for downstream effects. Finally, for general function classes (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),6, Q-MMR requires a minimax optimization, and the paper notes that practical implementations rely on no-regret oracles together with best responses over (s0,a0,r0,s1,a1,r1,…,sH,aH,rH),(s_0,a_0,r_0,s_1,a_1,r_1,\ldots,s_H,a_H,r_H),7, making efficient realization for large neural classes or RKHSs an implementation challenge (Li et al., 7 May 2026).

Within offline reinforcement learning, these points place Q-MMR at the intersection of theory-driven reweighting, value-function realizability, and coverage analysis. Its contribution is less a new heuristic estimator than a framework for understanding off-policy evaluation through recursive moment matching, with error terms that are explicitly decomposed across time, approximation, and reweighting variance (Li et al., 7 May 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Q-MMR.