Evaluate Personalization Across More Diverse Language Models

Investigate whether the human–LLM authorship gap and the limited effectiveness of inference-time personalization persist across language models with different parameter scales, architectures, and pretraining data.

Background

The study validates its findings using two 32B-scale, 4-bit-quantized model families: Qwen 3 and GLM-4. Both models therefore provide limited coverage of the broader language-model design space.

The authors explicitly identify evaluation on models differing in scale, architecture, and pretraining data as necessary to determine whether the observed authorship gap is a general property of inference-time personalization or is partly specific to the evaluated models.

References

We validate the main finding on two model families (Qwen~3 32B and GLM-4 32B), but both are 32B-scale 4-bit quantized models. More diverse generators---different parameter scales, architectures, and pretraining data---remain to be tested.

PersonalBench: Measuring the Authorship Gap in LLM Personalization  (2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Two generators, limited diversity”

PersonalBench evaluates blog-style writing from the early 2000s. Generalization to modern social media, email, or academic writing is untested.

PersonalBench: Measuring the Authorship Gap in LLM Personalization  (2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Single domain”