Evaluate Personalization Across More Diverse Language Models

Investigate whether the human–LLM authorship gap and the limited effectiveness of inference-time personalization persist across language models with different parameter scales, architectures, and pretraining data.

Background

The study validates its findings using two 32B-scale, 4-bit-quantized model families: Qwen 3 and GLM-4. Both models therefore provide limited coverage of the broader language-model design space.

The authors explicitly identify evaluation on models differing in scale, architecture, and pretraining data as necessary to determine whether the observed authorship gap is a general property of inference-time personalization or is partly specific to the evaluated models.

References

Whether these conclusions transfer to a low-resource, culturally distant language is, to our knowledge, untested.

— Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu  (2609.10758 - Adeeba et al., 9 Sep 2026) in Section 1, Introduction

We validate the main finding on two model families (Qwen~3 32B and GLM-4 32B), but both are 32B-scale 4-bit quantized models. More diverse generators---different parameter scales, architectures, and pretraining data---remain to be tested.

— PersonalBench: Measuring the Authorship Gap in LLM Personalization  (2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Two generators, limited diversity”

PersonalBench evaluates blog-style writing from the early 2000s. Generalization to modern social media, email, or academic writing is untested.

— PersonalBench: Measuring the Authorship Gap in LLM Personalization  (2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “Single domain”

Finally, robustness evaluation is necessarily bounded by currently available benchmarks and attack strategies. AI text generation is advancing rapidly, including the emergence of post-processing humanisation filters that deliberately reduce the statistical detectability of AI-generated text~\citep{Chakraborty24-Position}. A particularly challenging emerging scenario involves personally adapted AI systems that deliberately imitate an individual's writing style including their characteristic errors, vocabulary, and syntactic patterns making stylometric detection ineffective by design. Addressing such threats will require author-adaptive detection approaches and continued benchmark refreshment as both human writing and AI generation capabilities evolve.

— DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text  (2610.00883 - Mady et al., 1 Oct 2026) in Limitations, final paragraph; Appendix, subsection Detailed Dataset and Benchmarks Descriptions, paragraph Data currency note

Whether the trade-off structure is a property of style conditioning or of this configuration remains open. Replication across model scales, families, domains, and other languages is needed before generalizing.

— Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer  (2610.03163 - Popp et al., 2 Oct 2026) in Section “Limitations”