Deployment Failure Under Logged Reward Diagnostics

Determine whether a deployed recommendation policy would fail when the available logged match-and-score diagnostic shows only small changes and provides no positive evidence of reward improvement.

Background

The paper evaluates reward changes only for recommendations that match actions present in the test log. Because unobserved recommendations have no recorded ratings, the match-and-score diagnostic cannot estimate the rewards of a deployed policy or remove the selection bias inherent in logged data.

The authors explicitly state that the observed small changes are insufficient to determine whether a deployed policy would fail. The unresolved issue is therefore whether conclusions from this logged diagnostic correspond to actual deployment behavior, particularly when exposure effects, popularity, sparse items, and catalog-specific confounds are present.

References

These small changes cannot determine whether a deployed policy would fail.

Auditing Return Conditioning as a Control Knob: An Offline Diagnostic for Decision Transformer Recommendation  (2608.24815 - Wang, 25 Aug 2026) in Section 4.2, Scope and Failure Modes