Efficient and scalable learning of variable-length semantic IDs
Develop learning methods for variable-length semantic identifiers in recommender systems that are both efficient in representation and scalable to large-scale recommendation problems, ensuring practicality at industrial catalog sizes.
References
As a result, it remains unclear how to learn variable-length semantic identifiers that are both efficient and scalable to large-scale recommendation problems.
Several directions remain open. First, the current implementation relies on an offline diagnostic pass, and more efficient reachability estimation may further reduce post-training cost. Second, our experiments focus on a fixed semantic index with short four-token identifiers; extending the analysis to longer or hierarchical identifiers may reveal different hint-budget trade-offs.