Coat history-only ridge versus Phi reader

Determine whether the history-only ridge reader or the full-history Phi reader provides superior rating-prediction accuracy on the Coat dataset under matched historical-rating and item-metadata inputs.

Background

The paper compares a history-only ridge model with language-model readers using the same user histories and item metadata. On Coat, the ridge reader achieves a mean absolute error of 0.914, whereas Phi with the full history achieves 0.910. The resulting difference is extremely small, and its interval includes zero, so the evaluation does not establish superiority for either system or equivalence between them.

Because this comparison remains unresolved, determining whether either reader is genuinely better on Coat requires additional evidence, such as a larger evaluation or replication under the same matched-information design.

References

A history-only ridge reader outperforms Qwen in both domains and Phi on MovieLens, while the Coat Phi comparison is unresolved.

— A Controlled Audit of Personal AI Memory for Rating Prediction  (2610.02764 - Gupta, 2 Oct 2026) in Abstract; Section 4.3, “Reader choice and metric choice change the conclusion”