Policy-Aware Matrix Completion
- The paper introduces a bi-level framework that uses low-rank matrix completion from policy-induced sparse rewards to enhance learning efficiency.
- It integrates a confidence-weighted reward mechanism into policy optimization, reporting sample efficiency improvements of up to 2.1× over baseline methods.
- The approach is supported by theoretical guarantees including concentration inequalities and phase transition results, clarifying conditions for successful structural reward recovery.
Policy-Aware Matrix Completion (PAMC) is a framework for sparse-reward reinforcement learning that treats the reward function as a partially observed structured matrix and asks whether that matrix can be completed from the reward observations induced by the current policy. In its canonical formulation, the reward table is assumed to be low-rank or approximately low-rank, the observed entries are generated by the policy’s visitation distribution rather than by i.i.d. random sampling, and the completed reward is fed back into policy optimization through a confidence-weighted mechanism. The framework was introduced as part of a broader “structural reward learning” agenda that argues for a transition from essentially exponential behavior in the unstructured case to polynomial sample complexity under exploitable low-rank reward structure (Shihab et al., 4 Sep 2025).
1. Conceptual scope and defining idea
PAMC departs from the standard sparse-reward RL view that the problem is purely one of exploration. Its central claim is that many practical reward functions exhibit low-dimensional latent organization, so the full reward matrix may be recoverable from a sparse observation pattern if the completion procedure explicitly accounts for the fact that sampling is policy-dependent rather than exogenous. In this sense, the “policy-aware” qualifier is not ancillary: the observed reward entries are generated by the current policy, so the missing-data pattern is endogenous to the learning process (Shihab et al., 4 Sep 2025).
The framework is bi-level. The inner loop performs reward recovery via low-rank matrix completion from sparse observed rewards, and the outer loop updates the control policy using a confidence-weighted completed reward signal. The resulting workflow links matrix completion, reward-free representation learning, uncertainty calibration, and policy optimization into a single sequential procedure.
A common misconception is that PAMC is simply ordinary matrix completion applied inside RL. The distinction is sharper. Standard matrix completion theory typically assumes random missingness or uniform sampling, whereas PAMC models adaptive, sequential, policy-induced sampling. A second misconception is that PAMC presumes rewards are arbitrary tables that can always be completed if enough entries are seen. The motivating theorem is the opposite: without exploitable structure, sparse-reward observation is information-theoretically intractable in large spaces, which is why the low-rank structural hypothesis is central rather than optional (Shihab et al., 4 Sep 2025).
2. Mathematical formulation
The basic object is a reward matrix
whose entry is the reward for taking action in state . Only a sparse subset of entries is observed: Unlike classical completion settings, is produced by trajectories sampled under the current policy rather than by uniform random entry sampling (Shihab et al., 4 Sep 2025).
Policy dependence is summarized through the state-action distribution induced by the policy at time , and the exploration coefficient
0
This quantity measures how uniformly the policy sequence covers the reward matrix. If 1 is too small, some entries are effectively never seen; if it is sufficiently large, completion error can be controlled.
The completion model is factorized: 2 with state embeddings 3 and action embeddings 4. In matrix notation, the model is regularized by Frobenius norms and augmented with a confidence-calibration term: 5
The underlying structural assumption is that 6 has rank 7, or is approximately rank 8. The approximate model is written as
9
where 0 is rank 1 and 2 is a residual or misspecification term. Deviation from exact low rank is quantified by the tail singular value mass
3
The paper also uses an effective-rank notion,
4
as an empirical criterion for whether a reward matrix is structurally recoverable. For learned representations, PAMC introduces maps
5
so that
6
with 7 often a dot product or small network, and with an embedding-dimension requirement of the form
8
3. Structural assumptions and theoretical claims
The theoretical motivation for PAMC begins with an impossibility result. In episodic finite MDPs where rewards are observed independently with very low probability
9
the paper states that for any algorithm with sublinear regret there exists a reward family requiring sample complexity
0
to distinguish reward functions whose optimal values differ by 1. The proof idea is a Yao-style indistinguishability argument using reward matrices with a single active entry, so a learner that never visits and observes that specific state-action pair cannot determine which reward function it faces (Shihab et al., 4 Sep 2025).
Against this negative result, PAMC presents a positive low-rank regime. Under rank 2, the sample complexity becomes
3
which is polynomial when 4. This is presented as a phase transition: unstructured sparse rewards require essentially exhaustive discovery, whereas low-rank rewards can be recovered from a sample size that scales with intrinsic dimension rather than ambient size.
The bridge from completion accuracy to control performance is the regret-transfer bound
5
whenever
6
A related policy-aware completion bound states that if the reward matrix has rank 7 and embeddings of dimension 8 have distortion 9, then with
0
observations, the method achieves error 1 with probability at least 2, provided
3
Approximate low-rank structure is handled explicitly. The stated robust bound is
4
together with the simpler interpretation
5
The paper also gives a policy-dependent concentration inequality,
6
which makes the role of coverage explicit: better policy coverage improves completion accuracy (Shihab et al., 4 Sep 2025).
4. Algorithmic architecture
PAMC’s operational structure is a loop between acting, completing, calibrating, and updating. The agent collects trajectories 7 under the current policy 8, and these observations are added to the reward-training set. Independently or preemptively, the framework performs reward-free pretraining using contrastive predictive coding (CPC) on transition tuples 9, learning embeddings 0 from dynamics alone. The stated intuition is that if reward varies smoothly with respect to a dynamics-induced metric, then representations that preserve transition structure should also preserve reward-relevant structure (Shihab et al., 4 Sep 2025).
The smoothness condition is written as
1
Under this condition, CPC embeddings achieve distortion 2 using on the order of
3
transition samples.
Reward completion is updated every 4 episodes or when the policy changes enough, using ALS or similar methods. A held-out calibration set 5 is then used for conformal prediction. With nonconformity scores
6
the 7-quantile 8 defines the prediction interval
9
with finite-sample guarantee
0
The interval width is mapped to a confidence scalar, for example
1
Policy learning uses a confidence-weighted reward,
2
or an optimism or safe fallback reward, and updates 3 with PPO. The paper’s pseudocode is a bi-level loop: initialize policy and CPC embeddings, split observed rewards into training and calibration sets, collect trajectories, periodically update the factorization via ALS, compute 4, update confidence 5, compute confidence-weighted returns 6, update the policy with PPO, and repeat until convergence. Policy awareness enters twice: policy-induced sampling determines which matrix entries are observed, and the completed matrix influences the next policy update through the confidence-weighted reward signal (Shihab et al., 4 Sep 2025).
5. Genealogy in matrix completion and online decision-making
Although PAMC is introduced in sparse-reward RL, it sits within a broader line of work that treats observation selection as part of the problem rather than as an external sampling assumption. In “Matrix Completion With Selective Sampling,” the observation set 7 is deliberately designed when a subset of columns 8 is known to satisfy
9
That paper imposes structural constraints of the form
0
inside nuclear-norm completion and develops an explicit selective-sampling policy based on identifying invertible 1 submatrices. It also states that perfect reconstruction of 2 is achievable with
3
observations, described as necessary and sufficient under its setup. This earlier work does not formulate RL, but it supplies a direct precursor to PAMC’s policy-design intuition: spend sampling budget where structure is known, infer basis relations, and augment the completion objective with policy-derived constraints (Parkinson et al., 2019).
A second neighboring line is the matrix completion bandit formulation of online decision-making without informative covariates. There, each arm has its own low-rank reward matrix, action selection is carried out by an 4-greedy policy, and estimation proceeds via online gradient descent on a factorized model with inverse propensity weighting. Because only one noisy entry from one arm’s matrix is observed at each time, the data are adaptively collected and not i.i.d. This parallels PAMC’s rejection of fixed random missingness, but the modeling target differs: matrix completion bandits estimate arm-specific matrices for collaborative filtering policies, whereas PAMC completes a reward matrix coupled to policy-induced state-action occupancy and confidence-calibrated RL updates (Duan et al., 2024).
A separate source of potential confusion is optimization terminology. “Convergence of the majorized PAM method with subspace correction for low-rank composite factorization model” studies a majorized proximal alternating minimization method with subspace correction for generic low-rank composite factorization models. It is explicitly not a policy-aware matrix completion paper in the sense of exposure-aware or policy-constrained completion, but it is relevant at the solver level because PAMC-type models are typically solved via factorized nonconvex optimization. A plausible implication is that such convergence-certified alternating schemes could serve as optimization templates for factorized PAMC objectives, even though the paper itself does not introduce policy-aware sampling or reward-learning semantics (Tao et al., 2024).
6. Empirical profile, applications, and limitations
The empirical case for PAMC has two layers. First, the paper reports that across 25 domains there is a clear split: structured domains with low effective rank benefit substantially, borderline domains show modest gains, and unstructured domains can fail, as expected. Second, in a pre-registered sweep over 100 applications across sectors, exploitable structure is reported in about 52% overall, with higher rates in healthcare and robotics (Shihab et al., 4 Sep 2025).
In structured domains, PAMC is reported to achieve about 5 sample-efficiency improvement over exploration methods, 6 over structured baselines, and 7 over representation-learning baselines. The abstract summarizes this more generally as improving sample efficiency by factors between 8 and 9 compared to strong exploration, structured, and representation-learning baselines, while adding only about 20 percent computational overhead. The paper emphasizes that confidence calibration prevents catastrophic exploitation in low-structure or shifted regimes because the method abstains rather than forcing possibly wrong completions.
The intended application profile is therefore narrow in a principled sense. PAMC is especially attractive when samples are more expensive than compute, when reward matrices are approximately low-rank, and when safe abstention is preferable to blind exploitation. The paper is explicit that PAMC is not suitable for highly unstructured or adversarial rewards, that performance degrades when the reward matrix has high effective rank, that embedding learning may fail in very high-dimensional continuous control, and that poor calibration is the most dangerous failure mode. This suggests that PAMC should be understood not as a universal solution to sparse-reward RL, but as a structurally conditional method whose advantages arise precisely when reward recovery is a well-posed matrix completion problem under adaptive, policy-dependent sampling (Shihab et al., 4 Sep 2025).