- The paper introduces a unified reward learning framework using rigorous offline screening of billions of logs to optimize long-term user engagement.
- It integrates deep session engagement, negative rewards, and diversifying use case adoption into large-scale ranking models for improved personalization.
- Empirical results on Pinterest show consistent gains in session success rates, user retention, and engagement time across multiple product surfaces.
Model-Agnostic Downstream Rewards for Long-Term User Engagement
Introduction
Optimizing recommender systems for long-term user engagement and retention is a central challenge as the field moves beyond short-horizon behavioral signals. This work introduces a unified, model-agnostic framework for downstream reward (DR) learning to directly drive persistent user value in large-scale personalization systems. The approach is deployed and empirically validated at Pinterest across multiple product surfaces. The key contributions include robust reward signal selection via offline screening, production-grade reward derivation infrastructure, and successful transfer across heterogeneous ranking models in industry settings.

Figure 1: An illustration of user action patterns in P2P rabbit hole with closeup action at Pinterest.
Data-Driven Surrogate Rewards for Long-Term Retention
Direct retention optimization in recommender systems is fundamentally impeded by sparse, delayed return signals and ambiguous attribution. This work overcomes these challenges with a rigorous pipeline for surrogate downstream reward discovery. Offline analysis leverages billions of user activity logs to identify early-session behaviors both predictive of revisitation and denser than retention outcomes. The screening process normalizes features relative to each user’s historical baseline, leverages joint-feature models, and measures incremental predictive value.
Key findings include:
- Deep P2P exploration, saves, and diversified deep engagement are the most robust leading indicators of future retention.
- Shallow, high-volume behaviors—such as rapid browsing or broad but superficial category exposure—are misleadingly positive under naïve normalizations but become negative when controlling for exposure.
The pivot-day modeling, session-level Markov analysis, and revisitation latency metrics provide convergent validation: deeper actions consistently drive longer sessions, which in turn predict faster platform revisits. This motivates designing DR signals to promote depth, high-intent actions, and genuine exploration rather than superficial engagement.

Figure 2: Median hours to next revisit for sessions above versus below each duration threshold. Longer sessions consistently revisit sooner.


Figure 3: Probability of session exits of different closeup durations.
Reward Family Definitions and Infrastructure
Three complementary DR signal families are formulated:
- Deeper Session Engagement: Aggregates discounted cumulative engagement (e.g., saves, downloads, screenshots) following a recommendation, capturing sustained user value along multi-step trajectories.
- Negative Rewards: Penalizes low-quality actions such as shallow closeups that are rapidly abandoned without high-intent followup, with state-specific dwell-time thresholds determined by offline user-segment analysis.
- Use Case Adoption: Rewards engagement with items outside of a user’s entrenched interest clusters, operationalized as low cosine similarity to historical embeddings.
This structure enables flexible reward composition and aligns surrogate signals with long-horizon objectives. Importantly, negative signals (e.g., shallow closeups) are shown to have high action coverage and strong predictive utility when thresholded appropriately by user behavior clusters.
From a systems viewpoint, reward derivation is a central bottleneck at industrial scale. The initial DRv1 design relied on Spark-based pre-aggregation and static label tables—a process ill-suited to rapid iteration and combinatorial reward variants. The presented DRv2 infrastructure employs aligned per-user daily event sequences, with Ray-powered on-demand reward computation during training-data loading. This is a critical engineering advancement, yielding an order-of-magnitude reduction in experimentation timelines and enabling practical scaling of complex DR definitions.

Figure 4: DRv2 table structure illustration: per-user daily event sequences with aligned attribute arrays.
Integration in Ranking Models and Serving
Downstream rewards are incorporated as multi-head auxiliary objectives in large-scale ranking models (e.g., Pinnability), with tunable weightings for each head. Weight optimization is decoupled from model training and performed using hyperparameter search for product metric neutrality and retention lifts. Negative reward heads are implemented with negative utility, and reward integration extends seamlessly to all platform surfaces.
Empirical Results in Production Systems
A/B experiments on Pinterest demonstrate consistent, statistically significant improvements in both short-term meaningful engagement and retention-related metrics across Homefeed, Related Pins, Search, and Notifications.
Notable results:
- Deeper session engagement rewards produce up to +0.48% Successful Sessions (SS) and +0.10% total time spent, with pronounced lifts for non-core users.
- State-specific negative rewards yield +0.16% SS, -0.40% unsuccessful sessions, and up to +0.35% time spent, with clear improvements across DAU, WAU, and reduced hide/report rates. Tuning the penalty threshold by user segment is essential for positive cross-population effects.
- Use case adoption rewards increase SS (+0.10%), time spent (+0.15%), and drives users to engage with a broader array of interests, evidenced by a +0.18% increase in weekly active users with two or more use cases.
- Extension of deeper-session DR heads to Search, Related Pins, and Notifications preserves lift, e.g., +0.25% in search fulfillment and +0.15% in session frequency. Negative rewards similarly generalize and retain efficacy in non-entry surfaces.
The cross-surface extension is enabled with little engineering overhead due to the model-agnostic and infrastructure-agnostic DR design, and shows that optimizing for downstream value jointly across all surfaces can systematically improve holistic user journeys, mitigating local-product optimization traps.
Compared to RL-based recommendation [wo2017returning, wang2022surrogate, zhang2022multi, xu2023optimizing], this framework avoids the high complexity of reward engineering, policy gradient instability, and surface-specific customization. The approach is entirely agnostic to underlying model class or serving context, and DR signal discovery is driven by empirical analysis rather than heuristics. Supervised proxy reward approaches [wang2022surrogate, chang2023twin, si2024twin] are extended with new, empirically validated signals specific to visual discovery contexts and demonstrated deployment on multi-billion user platforms.
Implications and Future Directions
The presented framework demonstrates that model-agnostic surrogate reward construction, when validated by rigorous user-level analysis and supported by high-throughput infrastructure, can systematically improve both immediate and long-term user metrics in real-world recommender systems.
Theoretically, it bridges between short-term engagement optimization and long-horizon RL-style objectives via dense, interpretable proxies. Practically, the approach reduces engineering debt, enables rapid DR exploration, and supports flexible policy balancing between retention and myopic engagement.
Future work will likely focus on session-level and cross-session DR definitions (e.g., multi-day revisitation metrics), tighter integration of reward modeling and ranking (jointly trained architectures, latent-state models), and causal analysis to further decompose the impact of reward interventions on true retention.
Conclusion
This work delivers a rigorous, scalable model-agnostic framework for long-term user engagement optimization in recommender systems. Surrogate downstream reward learning—anchored by robust offline screening, rapid label derivation, and demonstrated product impact—emerges as a practical and effective paradigm for aligning industrial recommendation with durable user value.