Papers
Topics
Authors
Recent
Search
2000 character limit reached

Creative Homogeneity in LLMs

Updated 1 May 2026
  • Creative Homogeneity in LLMs is the phenomenon where models generate outputs with reduced lexical, semantic, and structural diversity, leading to repetitive patterns.
  • Empirical metrics such as embedding similarity, self-BLEU scores, and entropy measures reveal a significant collapse in diversity across various creative and divergent tasks.
  • Mitigation strategies including prompt engineering, diversity-promoting training objectives, and human-in-the-loop audits are proposed to counteract systemic biases and enhance variability.

Creative homogeneity in LLMs denotes the systemic tendency of these models to generate outputs that are semantically, structurally, or stylistically similar across repeated samples, tasks, and even across distinct models. This phenomenon arises from the statistical, architectural, and alignment-centric foundations of current LLMs, and is evidenced by reduced lexical, conceptual, and perspectival diversity relative to human population baselines. Empirical studies confirm that creative homogeneity pervades everything from divergent‐thinking tests and narrative generation to open-ended ideation and social group portrayals, raising concerns for language diversity, cognitive pluralism, fairness, and the future trajectory of both AI-generated and human communication.

1. Formal Definitions and Core Metrics

Creative homogeneity is instantiated through several forms—semantic, lexical, structural—that are quantified via clustering, pairwise similarity, entropy, and divergence metrics:

  • Semantic Homogeneity: Let S={P1,,PK}S = \{\mathcal{P}_1, \dots, \mathcal{P}_K\} be KK model-generated texts for a fixed prompt. Embedding vectors e(Pk)Rde(\mathcal{P}_k)\in\mathbb{R}^d yield mean pairwise cosine similarity:

Sintra(S)=2K(K1)i<jcos(e(Pi),e(Pj))S_\text{intra}(S) = \frac{2}{K(K-1)}\sum_{i<j}\cos\left(e(\mathcal{P}_i), e(\mathcal{P}_j)\right)

Higher values manifest greater homogeneity; low diversity is equally evidenced by high self-BLEU and low Distinct-nn scores (Wenger et al., 31 Jan 2025, Jiang et al., 27 Oct 2025, Yun et al., 25 May 2025).

  • Population Variability: Measured by embedding-based spread or pairwise sentence similarity among all responses in a population (human or LLM). Wenger & Kenett compute (Wenger et al., 31 Jan 2025):

Vt(P)={1cos(S(Rip),S(Rjp))  ij,p}\mathcal{V}_t(\mathcal{P}) = \left\{1 - \cos\big(\mathcal{S}(\mathbf{R}_i^p), \mathcal{S}(\mathbf{R}_j^p)\big)\ |\ i \neq j, p\right\}

LLM–LLM distances are significantly lower than human–human on divergent-thinking tasks.

In multi-model or multi-sample contexts, inter-model and intra-model similarities are also central: S_intra (within model), S_inter (between models) (Jiang et al., 27 Oct 2025).

2. Empirical Manifestations and Scope

Creative homogeneity is robustly observed across models, tasks, and evaluation settings:

  • Population-Level Collapse: On divergent-thinking tests (Alternative Uses, Forward Flow, Divergent Association), LLM outputs are significantly more similar to each other (AUT: VAUTLLM=0.459\mathcal{V}_\text{AUT}^\text{LLM}=0.459, KK0) despite comparable individual originality scores (Wenger et al., 31 Jan 2025).
  • Instruction and Chat Tuning: Fully structured prompts (system/user/assistant tokens) anchor models to restrictive output spaces; “diversity collapse” persists even at high temperature, and removing chat tokens increases output diversity by 2x–3x (Yun et al., 25 May 2025).
  • Narrative Homogenization: In both English and Chinese story generation, LLMs exhibit fixed, repetitive plot paradigms with low narrative‐function entropy (KK1) and high Jensen-Shannon divergence from human distributions. Even high-capacity models like Deepseek and Qwen default to “hero-struggle-victory” templates (Ma et al., 15 Mar 2026).
  • Open-Ended Ideation: On the Infinity-Chat corpus, 79% of open-ended prompts yield S_intra > 0.8; human population variability dwarfs that of both intra- and inter-model LLM generations (Jiang et al., 27 Oct 2025).
  • Cross-Model Agreement: LLMs used as creative judges agree with each other at Spearman KK2 > 0.7 and even match expert-constructed oracles (KK3), reinforcing output homogeneity in both creation and assessment (Rabeyah et al., 2024).

3. Mechanistic Sources and Theoretical Frameworks

Multiple, orthogonal mechanisms underlie creative homogeneity:

  • Statistical Centrality and Averaging: Next-token prediction implicitly rewards frequent, high-probability continuations, pulling outputs toward “average” rhetorical, stylistic, or conceptual modes (“kitsch” metaphor) (Uhlmann, 20 Sep 2025).
  • Format-Induced Collapse: Instruction-finetuning on chat-style (multi-turn) templates instills fixed generation priors, rapidly restricting output variability after just a few decoding steps (Yun et al., 25 May 2025).
  • Alignment and Uncertainty Minimization: Alignment via RLHF or SFT steers outputs toward low-uncertainty, “safe” continuations, suppressing constructive ambiguity and leading to bland, risk-free narratives (Sui, 18 Feb 2026).
  • Cognitive Process Compression: Without explicit divergent–convergent phase separation, LLMs conflate idea generation with immediate constraint satisfaction, pruning uncommon ideas and drift toward a hive-mind (Nguyen et al., 29 Dec 2025).
  • Training Data and Social Bias: Differential coverage and stereotype alignment in corpora cause subordinate social groups to be depicted with markedly higher narrative homogeneity, mirroring the human out-group homogeneity effect (Lee et al., 2024).
  • Reward Models and Model Judges: LM judges and reward models, trained to replicate aggregate “quality,” demonstrate strong cross-model calibration but fail to account for the pluralism apparent in human preferences (Jiang et al., 27 Oct 2025, Rabeyah et al., 2024).

4. Social and Cultural Implications

The homogenizing tendencies of LLMs propagate beyond algorithmic evaluations into broader social and epistemic contexts:

  • Linguistic and Cultural Flattening: Lexical, syntactic, and stylistic diversity in LLM-assisted texts is systematically lower (TTR: −15–20%, Shannon entropy suppressed by comparable margins) (Sourati et al., 2 Aug 2025, Uhlmann, 20 Sep 2025).
  • Perspective and Reasoning Narrowing: LLMs collapse perspectival spread (Jaccard similarity: 0.75 LLM vs 0.45 human), erasing rare beliefs or reasoning strategies (Sourati et al., 2 Aug 2025).
  • Reinforcement Loops: Recursive use of LLM-generated content in communications and knowledge bases turns statistically dominant patterns into cultural norms, amplifying convergence and diminishing pluralism over time (Sourati et al., 2 Aug 2025).
  • Creative Erasure and Fairness: In domains such as narrative generation about socially subordinate groups, homogeneity compounds social bias, narrowing both individuating detail and trait diversity (Lee et al., 2024).
  • Educational and Ideational Impact: Homogeneous LLM-generated educational materials propagate redundant or repetitive problem structures, impinging on divergent thinking and novelty for learners (Nguyen et al., 29 Dec 2025).

5. Mitigation Strategies and Interventions

Approaches for ameliorating creative homogeneity span data, model, and user-facing solutions:

  • Prompt and Decoding Engineering:
  • Architectural and Objective Modifications:
    • Training with diversity-promoting objectives or multi-persona/debate-driven architectures (Sourati et al., 2 Aug 2025).
    • Fine-tuning with anti-causal and structural-contrastive losses; explicit narrative-function supervision for fiction domains (Ma et al., 15 Mar 2026).
    • Uncertainty-aware alignment (ambiguity-permissive reward models, credal-set planning) to re-inject productive unpredictability (Sui, 18 Feb 2026).
  • Data and Evaluation Practices:
  • Critical and Regulatory Approaches:
    • “Kitsch studies” and critical rhetoric frameworks to educate users and guide regulatory audits for style and content diversity (Uhlmann, 20 Sep 2025).
    • Diversity-focused curriculum and alignment testing within foundation model development (Jiang et al., 27 Oct 2025).

6. Persistent Challenges and Open Problems

Despite methodological innovations and some gains, significant obstacles remain:

  • Superficial vs. Deep Diversity: Many interventions (e.g., temperature scaling, anti-kitsch prompting) boost superficial variation without restoring deep cultural or reasoning heterogeneity (Sourati et al., 2 Aug 2025).
  • Scalability of Diversity Curation: Ensuring pluralistic representations at web scale requires scalable data collection and labeling solutions.
  • Evaluation in Unscored Domains: Formal measurement of creativity and diversity lags in poetry, humor, argumentation, and long-form storytelling. LM-based judges, though consistent, cannot presently capture the pluralism observable in human populations (Jiang et al., 27 Oct 2025, Rabeyah et al., 2024).
  • Societal Feedback Loops: As LLMs proliferate as both content generators and evaluators, their structural bias towards consensus and similarity risks accelerating epistemic and cultural convergence beyond language modeling per se (Sourati et al., 2 Aug 2025, Jiang et al., 27 Oct 2025).

7. Summary Table: Major Sources, Metrics, and Empirical Findings

Paper (arXiv ID) Core Metric(s) Primary Homogeneity Finding
(Wenger et al., 31 Jan 2025) KK4 (embed dist) LLM–LLM similarity KK5 Human–Human, despite equal originality
(Jiang et al., 27 Oct 2025) S_intra, S_inter (embed sim) Intra- and inter-model S > 0.8 on 50+% of open-ended tasks
(Yun et al., 25 May 2025) D_sem, D_topic, self-BLEU Format induces 2–3x collapse in diversity, persists at T=1.0
(Sui, 18 Feb 2026) PPL, NLL, PMI, CPMI LLM uncertainty 2–8x lower than humans, correlates to triteness
(Nguyen et al., 29 Dec 2025) SemDiv, LexDiv, Vendi Score Two-phase prompts (CreativeDC) increase effective distinctivity
(Lee et al., 2024) Mean standardized embedding Subordinate groups are more homogeneous (race, gender effects)
(Ma et al., 15 Mar 2026) JSD, BERTScore, Entropy Narrative function collapse, fixed high-frequency plot templates
(Sourati et al., 2 Aug 2025) TTR, entropy, Jaccard, KL Lexical, syntactic, and reasoning diversity all shrink in LLMs
(Rabeyah et al., 2024) Spearman KK6 (AUT judge) LLM–LLM KK7 > 0.7 for creativity evaluation

The creative homogeneity of LLMs is a multidimensional, system-level phenomenon resulting from the statistical, algorithmic, and socio-cultural processes governing current models and their deployment. It is grounded in robust empirical findings, has measurable expression across creative and non-creative tasks, and presents a suite of technical and societal challenges that require algorithmic, evaluative, and cultural mitigation strategies (Wenger et al., 31 Jan 2025, Sourati et al., 2 Aug 2025, Sui, 18 Feb 2026, Ma et al., 15 Mar 2026, Jiang et al., 27 Oct 2025, Yun et al., 25 May 2025, Rabeyah et al., 2024, Lee et al., 2024, Nguyen et al., 29 Dec 2025, Uhlmann, 20 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Creative Homogeneity in LLMs.