Papers
Topics
Authors
Recent
Search
2000 character limit reached

Less is More: Benchmarking LLM Based Recommendation Agents

Published 28 Jan 2026 in cs.IR and cs.LG | (2601.20316v1)

Abstract: LLMs are increasingly deployed for personalized product recommendations, with practitioners commonly assuming that longer user purchase histories lead to better predictions. We challenge this assumption through a systematic benchmark of four state of the art LLMs GPT-4o-mini, DeepSeek-V3, Qwen2.5-72B, and Gemini 2.5 Flash across context lengths ranging from 5 to 50 items using the REGEN dataset. Surprisingly, our experiments with 50 users in a within subject design reveal no significant quality improvement with increased context length. Quality scores remain flat across all conditions (0.17--0.23). Our findings have significant practical implications: practitioners can reduce inference costs by approximately 88\% by using context (5--10 items) instead of longer histories (50 items), without sacrificing recommendation quality. We also analyze latency patterns across providers and find model specific behaviors that inform deployment decisions. This work challenges the existing ``more context is better'' paradigm and provides actionable guidelines for cost effective LLM based recommendation systems.

Summary

  • The paper tests the hypothesis that longer user purchase histories improve LLM-based product recommendations using four production LLMs and finds no statistically significant improvement in recommendation quality with increased context length.
  • The study uses a composite quality score that combines keyword overlap and category match to evaluate recommendation performance and find average quality scores ranging from 0.16 to 0.23 across all model-context combinations.
  • The results of the study offer a potential deployment guideline to significantly reduce costs by using a minimal sufficient context (5-10 items), which achieves quality results at approximately one-eight the token expense.

Overview and motivation

This paper presents a controlled benchmark examining whether longer user purchase histories improve the quality of LLM-based product recommendations. The authors, Chauhan and Venkateswarlu (UC Santa Cruz and Georgia Tech), test a widely held practitioner assumption—that more context yields better personalization—by varying history length from 5 to 50 items across four production LLMs: GPT-4o-mini (OpenAI), DeepSeek-V3, Qwen2.5-72B (both via Together AI), and Gemini 2.5 Flash (Google). The central research question is whether providing more user purchase history improves LLM-based recommendation quality.

The motivation is both scientific and economic. LLM-based recommenders process textual purchase histories directly rather than learned embeddings, and agentic recommendation frameworks such as RecMind and Agent4Rec treat context as a finite resource shared across planning, retrieval, and generation. If extended histories do not improve quality, then token-based API pricing makes long contexts a pure cost liability. The paper's headline claim is that they do not: quality is statistically flat across all tested context lengths, while token consumption grows roughly 8×.

Experimental design

The study uses the REGEN dataset (Reviews Enhanced with GEnerative Narratives), which extends Amazon Product Reviews with item metadata, review information, ground-truth next purchases, and explanatory narratives. The authors restrict to the Office Products category and select users with at least 51 interactions (50 for context plus one held-out target), yielding 189 eligible users from 89,489 total; 50 are sampled with a fixed seed for reproducibility.

A within-subject design evaluates the same 50 users at five context lengths (5, 10, 15, 25, 50 items), always using the most recent kk items to preserve recency. Each model receives a standardized prompt listing item titles, categories, and ratings, and is asked to predict the next purchase with reasoning. Three metrics are collected:

  • Quality: a composite score, 0.7×KeywordScore+0.3×CategoryMatch0.7 \times \text{KeywordScore} + 0.3 \times \text{CategoryMatch}, where KeywordScore is Jaccard-like keyword overlap with the ground truth and CategoryMatch is binary.
  • Latency: wall-clock time from request initiation to response receipt.
  • Token count: input tokens as reported by each API, as a proxy for cost.

The composite metric follows evaluation practices in prior zero-shot recommendation work where lexical similarity is weighted above coarse category matching, though the authors acknowledge it does not capture ranking-aware utility.

Main results: flat quality curves

The core finding is that recommendation quality does not improve with context length. Quality scores range from 0.16 to 0.23 across all 20 model–context combinations, with overlapping confidence intervals throughout:

Model Ctx=5 Ctx=10 Ctx=15 Ctx=25 Ctx=50 Δ (5→50)
DeepSeek-V3 0.21 ± 0.10 0.23 ± 0.10 0.22 ± 0.10 0.20 ± 0.11 0.20 ± 0.11 −0.01
Qwen2.5-72B 0.20 ± 0.10 0.21 ± 0.11 0.19 ± 0.08 0.19 ± 0.09 0.18 ± 0.10 −0.02
GPT-4o-mini 0.17 ± 0.08 0.20 ± 0.08 0.19 ± 0.09 0.19 ± 0.08 0.19 ± 0.09 +0.02
Gemini 2.5 Flash 0.18 ± 0.10 0.21 ± 0.13 0.19 ± 0.10 0.19 ± 0.10 0.16 ± 0.09 −0.02
Average 0.19 0.21 0.20 0.19 0.18 −0.01

Statistical analysis supports the null result: paired t-tests comparing context 5 versus 50 yield p>0.05p > 0.05 for every model, and a repeated measures ANOVA across all five lengths shows no significant main effect of context (F(4,196)=1.12F(4,196)=1.12, p=0.35p=0.35). The average quality change across models is −0.01—essentially zero.

Two aspects of this result deserve emphasis. First, its universality: the flat pattern holds across four providers with different architectures and training regimes, which the authors interpret as evidence of a fundamental limitation in how current LLMs utilize extended user histories rather than a model-specific artifact. Second, its practical consequence: because token usage rises from an average of 288 tokens at context 5 to 2,371 tokens at context 50 (an 8.2× increase), practitioners can cut inference costs by approximately 88% by truncating histories to 5–10 items with no measurable quality loss. For a system issuing one million API calls daily, the authors estimate annual savings of $300,000 or more.

Latency behavior

Latency patterns are strongly model-specific rather than uniform. Qwen2.5-72B maintains stable latency of 4.11–4.39 s regardless of context length, suggesting that network round-trip time and API overhead dominate over token processing below ~3,000 input tokens. GPT-4o-mini is moderate (4.54–5.86 s); DeepSeek-V3 is variable (6.38–10.33 s); and Gemini 2.5 Flash exhibits monotonically increasing latency with context, from 9.97 s at 5 items to 15.44 s at 50 items. These profiles inform deployment choices: Qwen2.5-72B suits latency-sensitive real-time serving, while GPT-4o-mini offers reasonable cost-efficiency at short contexts. Notably, Gemini's scaling behavior means longer contexts incur both cost and responsiveness penalties on that provider.

Interpretation

The authors offer four candidate explanations for why additional context fails to help. The "Lost in the Middle" phenomenon suggests LLMs underutilize mid-context information, so most of a 50-item history falls into an ignored region. Recency bias implies the most recent few purchases capture current interests as well as longer windows do, consistent with sequential recommendation findings. Signal saturation posits that the first several items establish preferences, with later items contributing noise. Finally, task difficulty imposes a ceiling: predicting the exact next product from free-text descriptions is inherently hard, so even perfect context utilization may not raise quality substantially. These explanations are offered as plausible mechanisms rather than experimentally isolated causes—the paper does not disentangle them empirically.

Limitations

The paper concedes several constraints on its conclusions. Evaluation covers only the Office Products domain of REGEN; results may differ in domains such as fashion or entertainment where preference dynamics differ. The composite quality metric captures semantic similarity but not real-world recommendation utility, and no human evaluation was performed; ranking metrics such as MRR and NDCG are deferred to future work. Only four general-purpose models were tested—fine-tuned or recommendation-specialized models might behave differently. Absolute quality scores (0.17–0.23) are modest, reflecting task difficulty, so the contribution rests on relative patterns rather than absolute performance. Finally, a single simple prompt template was used; more sophisticated prompting strategies could alter how context length affects quality. Sample size is also modest (50 users), though the within-subject design mitigates user-level variance.

Conclusion

This benchmark demonstrates that increasing user purchase history from 5 to 50 items produces no statistically significant improvement in LLM recommendation quality across four state-of-the-art models, while inflating token costs by approximately 8×. The consistency of the null effect across providers suggests a structural property of current LLMs' use of extended user histories, and the associated cost analysis yields a concrete deployment guideline: minimal sufficient context of 5–10 items preserves quality at roughly one-eighth the token expense. Open questions remain regarding domain generality, ranking-aware evaluation, human-judged quality, and whether prompt engineering or adaptive context selection can change the context-length dynamics established here.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.