Evaluate Personalization Across More Diverse Language Models
Investigate whether the human–LLM authorship gap and the limited effectiveness of inference-time personalization persist across language models with different parameter scales, architectures, and pretraining data.
References
We validate the main finding on two model families (Qwen~3 32B and GLM-4 32B), but both are 32B-scale 4-bit quantized models. More diverse generators---different parameter scales, architectures, and pretraining data---remain to be tested.
— PersonalBench: Measuring the Authorship Gap in LLM Personalization
(2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Two generators, limited diversity”
PersonalBench evaluates blog-style writing from the early 2000s. Generalization to modern social media, email, or academic writing is untested.
— PersonalBench: Measuring the Authorship Gap in LLM Personalization
(2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Single domain”