Evaluate Personalization Across More Diverse Language Models
Investigate whether the human–LLM authorship gap and the limited effectiveness of inference-time personalization persist across language models with different parameter scales, architectures, and pretraining data.
References
Whether these conclusions transfer to a low-resource, culturally distant language is, to our knowledge, untested.
We validate the main finding on two model families (Qwen~3 32B and GLM-4 32B), but both are 32B-scale 4-bit quantized models. More diverse generators---different parameter scales, architectures, and pretraining data---remain to be tested.
PersonalBench evaluates blog-style writing from the early 2000s. Generalization to modern social media, email, or academic writing is untested.
Finally, robustness evaluation is necessarily bounded by currently available benchmarks and attack strategies. AI text generation is advancing rapidly, including the emergence of post-processing humanisation filters that deliberately reduce the statistical detectability of AI-generated text~\citep{Chakraborty24-Position}. A particularly challenging emerging scenario involves personally adapted AI systems that deliberately imitate an individual's writing style including their characteristic errors, vocabulary, and syntactic patterns making stylometric detection ineffective by design. Addressing such threats will require author-adaptive detection approaches and continued benchmark refreshment as both human writing and AI generation capabilities evolve.
Whether the trade-off structure is a property of style conditioning or of this configuration remains open. Replication across model scales, families, domains, and other languages is needed before generalizing.