Papers
Topics
Authors
Recent
Search
2000 character limit reached

From Raw IDs to Semantic Planning: How Recommender Systems Utilize Information at Scale

Published 10 Jul 2026 in cs.IR | (2607.09540v1)

Abstract: The evolution of recommender systems can be explored by asking how they utilize information at scale. Throughout most of the historical period under consideration during the past two decades, industrial systems have relied on raw IDs, which are discrete, globally unique, and semantically opaque identifiers that enable exact lookup, logging, and item-specific memorization at scale. Over time, however, recommender systems have sought to utilize richer sources of information, including item content, context, multimodal signals, and cross-domain structure. This development has led to a new stage in which part of such information is no longer used solely as auxiliary features around item identity, but is increasingly encapsulated in semantic IDs that provide a more structured, model-facing form of identity. We argue that this shift goes beyond the rise of generative recommendation over traditional methods. Indeed, it reflects a broader evolution in how recommender systems utilize information under industrial-scale constraints. This paper looks at the past, present, and future to examine three connected questions: why raw IDs dominated the early development of recommender systems, why semantic information is increasingly being encapsulated in IDs today, and what may come next once recommendations move beyond semantic retrieval. In particular, we introduce semantic planning as a possible future direction in which the system first predicts the semantic target of the next exposure, and only then instantiates that target as a specific item or generated creative. We further argue that such a shift may require changes not only in model design but also in evaluation and in the way recommender systems coordinate the objectives of users, platforms, and providers.

Summary

  • The paper introduces semantic planning, a framework that integrates multi-stakeholder objectives with enriched semantic ID representations.
  • It details the transition from raw, opaque identifiers to structured, semantic IDs that enhance cold-start resilience and cross-domain transfer.
  • The study advocates new evaluation protocols and deployment strategies, aligning system decisions with diverse user, provider, and platform goals.

From Raw IDs to Semantic Planning: A Technical Analysis of Information Utilization in Recommender Systems

Introduction

The paper "From Raw IDs to Semantic Planning: How Recommender Systems Utilize Information at Scale" (2607.09540) examines the architectural and methodological progression of industrial recommender systems, focusing on how systems exploit various forms of information to support decision-making under industrial constraints. The analysis is presented through a historical lens, revealing a trajectory from the reliance on discrete, semantically opaque raw IDs, through hybridization with semantic enrichment, toward the emergence of structured semantic IDs, and postulating semantic planning as the next critical paradigm.

Historical Perspective: Operational Role of Raw IDs

Original large-scale recommenders predominantly utilized raw IDs—stable, unique identifiers employed as operational anchors throughout the entire serving pipeline. These IDs provided a robust mechanism for lookup, logging, and behavioral aggregation, supporting the stability and auditability required in high-throughput, multi-stakeholder environments. Raw IDs facilitated exact memorization of item-specific behavior and commercial value, but this design led to pronounced limitations regarding cold-start scenarios, scalability of embedding tables, and inefficiencies in representing cross-item semantic relationships. The resulting architecture often necessitated learning the nuances of every new item from scratch, with considerable redundancy for semantically similar entities.

Present Day: Towards Semantic IDs and Unified Representations

With advances in modeling, recommender systems began to integrate richer semantic signals—textual content, visual modalities, contextual cues, and behavioral sequences—as auxiliary features. However, the identity layer often remained predicated on raw IDs, maintaining the split between semantic understanding and operational identity. The introduction of semantic IDs marks a fundamental shift where the identifiers themselves become carriers of semantic structure, thus unifying content and behavioral signals within the ID space. Semantic IDs serve as discrete token sequences consumable across generative, retrieval-oriented, and ranking models, bypassing the need for cross-space alignment and facilitating transfer and generalization across domains and modalities.

This migration towards semantic IDs is enabling system-wide unifications: within-domain (enabling cold-start and structural generalization), cross-domain (enabling unified identity spaces for transfer), and inter-task (storming the boundary between search and recommendation). Yet, despite the sophistication in user-item matching, the multi-stakeholder problem—reasoning over conflicting objectives of users, providers, and platforms—remains largely unsolved under the current retrieval-centric framework.

Future Paradigm: Semantic Planning

The paper introduces the concept of semantic planning as the logical progression of recommender systems, superseding semantic retrieval with explicit, multi-stakeholder intent-driven recommendation. Instead of directly matching user state to catalog items, semantic planning inserts an intermediate layer that predicts a semantic target, capturing latent or explicit objectives of all participating stakeholders. This allows the recommender to instantiate the target through various exposures (items, offers, creatives), shifting the system's function from mere item selection to orchestrating exposures that align with well-articulated goals and constraints.

Semantic planning also elevates the system's ability to recognize and signal unmet semantic requirements—exposing limitations of the current catalog and informing platform and provider strategy. Additionally, it reflects a theoretical unification: planning over a joint space of user and ecosystem value, rather than mere myopic user-item preference modeling. This transition promises explicit control, enhanced adaptability to inventory dynamics, and principled multi-objective optimization.

Evaluation and Deployment Implications

The shift to semantic planning mandates a departure from traditional evaluation protocols predicated on fixed relevance judgments and hit-rate metrics. The paper argues for the development of evaluation agents and long-term goal-tracking mechanisms to assess the resonance of system decisions with user and ecosystem-level intent fulfillment. Key open challenges include the stability and lifecycle management of semantic IDs amidst inventory churn, the grounding of planning targets to ensure viability given current platform capabilities, and the partial observability of provider and platform objectives.

From a deployment standpoint, semantic planning facilitates more nuanced interfaces between ranking, advertising, and content systems, driving coordination via shared intermediate objectives. This enables platforms to graduate from passive traffic distribution to active demand aggregation and strategic provider feedback, fostering tighter supply-demand coupling and adaptive content strategies.

Practical and Theoretical Implications

Practically, the advancement toward semantic IDs and planning is poised to reduce cold-start brittleness, improve transfer across application domains, and enable more adaptive user and inventory management. Theoretically, it expands the scope of recommendation research from matching and ranking toward sequential, multi-objective planning, opening intersections with causal inference, agentic user simulation, and content generation. The anticipated hybrid architectures—maintaining raw IDs for operational integrity and semantic IDs for representational flex—are likely to be central to next-generation production systems.

Conclusion

The trajectory traced in the paper reflects an increasing sophistication in how recommender systems handle scale, heterogeneity, and stakeholder complexity. The move from raw IDs to semantic IDs represents a significant restructuring of the information substrate underpinning recommendations, setting the stage for efforts in semantic planning. Future breakthroughs will hinge on orchestrating semantic representation, planning objectives, and multi-stakeholder tradeoffs, with the ultimate test being the capacity to translate diverse ecosystem needs into coherent, purposeful, and actionable exposures. This paradigm, if realized, positions recommender systems as central agents in ecosystem coordination rather than mere engines of personalized retrieval.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 19 likes about this paper.