Cultural Neutrality in Theory and AI
- Cultural Neutrality is defined as unbiased cultural transmission, symmetric measurement, and an AI design goal that respects diverse cultural norms.
- Empirical studies reveal significant discrepancies in offensiveness judgments and model outputs across cultures, challenging the notion of universal neutrality.
- Methodologies including cultural value surveys, copula-graphical models, and neuron-level analyses highlight the need for contextual sensitivity and non-dominance in practice.
Cultural neutrality denotes, in different research traditions, either a null condition of unbiased cultural transmission, a comparative measurement principle that does not privilege one national value structure, or a design goal for sociotechnical systems that should not default to the norms of a dominant culture. Contemporary work repeatedly challenges strong versions of that ideal. Offensiveness judgments differ substantially across 21 countries and 8 geo-cultural regions, mainstream GPT models resemble English-speaking and Protestant European countries, multilingual evaluation exposes language-dependent fairness failures, and multilingual text-to-image systems often produce culturally neutral or English-biased outputs (Davani et al., 2023, Tao et al., 2023, Huang et al., 13 Jul 2025, Shi et al., 21 Nov 2025).
1. Definitions across cultural theory, measurement, and AI
The literature uses “cultural neutrality” in several non-equivalent senses. In cultural transmission theory, neutrality refers to the absence of systematic preference among variants. In the overlapping-generations model of cultural evolution, variants are treated as selectively equivalent and innovation introduces new variants at rate ; in the archaeological Wright–Fisher infinite-alleles setting, neutrality is defined as unbiased copying and is explicitly distinguished from drift, which is finite-population stochasticity rather than absence of bias (O'Dwyer et al., 2017, Madsen, 2012).
In comparative cultural measurement, neutrality is framed less as value-freedom than as symmetry of representation. “Cultures as networks of cultural traits” defines national culture as the combination of marginal distributions of shared cultural traits and the latent network structure among them, so that no country is treated as the benchmark and no single low-dimensional axis is assumed to exhaust cultural structure. Cultural distance is then measured by the Jeffreys’ divergence between country-specific copula graphical models, with separate components for trait distributions and network structure (Benedictis et al., 2020).
In AI research, the term is increasingly reformulated. Dialogue and safety work does not equate neutrality with removal of culture. Rather, a culturally neutral or culturally safe system is described as one that avoids culturally biased defaults, Western-centric assumptions, or hostile premises while remaining respectful, context-aware, and aligned with the user’s cultural setting. This suggests a shift from “culture-free” behavior to non-dominance, contextual sensitivity, and explicit accommodation of plural norms (Cao et al., 2024, Banerjee et al., 2024).
2. Neutrality as an empirical null in cultural transmission
Within cultural evolution, neutrality functions as a testable null hypothesis. “Inferring processes of cultural transmission” derives an analytical late-time progeny distribution under neutrality and shows that it has two regimes: a power-law phase with a universal exponent of , followed by an exponential cutoff. The paper’s central methodological claim is that rare variants are decisive: analyses restricted to abundant variants can make non-neutral processes appear neutral, whereas the full Australian baby-name distribution is better replicated by anti-novelty bias than by neutrality (O'Dwyer et al., 2017).
Archaeological applications complicate this null. “Unbiased Cultural Transmission in Time-Averaged Archaeological Assemblages” shows that time averaging systematically alters neutral-model predictions. Richness is inflated, evenness is flattened, and neutrality tests suffer serious Type I error under time averaging. The relevant timescale is determined by mean trait lifetime, so assemblage duration relative to trait lifetime governs when a genuinely neutral process will falsely appear non-neutral (Madsen, 2012).
Measurement work reaches a related conclusion by a different route. In the copula-graphical framework for national culture, the distance decomposition
formalizes the idea that cultural similarity cannot be reduced to average trait values alone. Empirically, the marginal component correlates strongly with the Inglehart–Welzel index, about $0.89$–$0.90$, whereas the network component is essentially uncorrelated with it, about . A plausible implication is that apparently “neutral” low-dimensional summaries can obscure culturally relevant dependence structure (Benedictis et al., 2020).
3. Moral judgment, moderation, and pluralistic safety
A major contemporary challenge to cultural neutrality arises in moderation and safety labeling. “Disentangling Perceptions of Offensiveness” argues that offensiveness is fundamentally subjective and moralized rather than a culturally neutral label. In a cross-cultural study with 4309 participants, 21 countries, and 8 geo-cultural regions, a one-way ANOVA with cross-classified mixed-level regression found significant regional effects, , and 25 of 28 region pairs differed significantly. These differences remained significant after controlling for age, gender, self-reported socio-economic status, and whether participants were given a definition of offensiveness (Davani et al., 2023).
The same study treats moral values as mediators rather than background noise. Significant mediation was found for Care and Purity, with average causal mediating effects of Care: and Purity: . The formal decomposition
separates country-level moral climate from individual deviation, and the results show that the individual-level deviation matters more than the country average: for Care, country-level 0 versus individual deviation 1; for Purity, country-level 2 versus individual deviation 3 (Davani et al., 2023).
Safety-alignment datasets show a parallel pattern. “Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment” meta-analyzes 8 safety datasets—DIVE, CulturalFrames, PRISM, DICES-990, NLPositionality, D3, CREHate, and Severity—and finds that cultural zone membership explains variance in safety ratings beyond demographics in 6 out of 8 datasets, with reported significance at 4 after Benjamini–Hochberg correction. The paper further estimates that about 10.53% of items are culturally sensitive, meaning that they are likely to be misclassified as safe without adequate cultural representation. LLMs do not reliably emulate quadrant-level judgments as rater surrogates, though they can help triage culturally sensitive items for human review (Saakyan et al., 29 May 2026).
Taken together, these results undermine the assumption that moderation or safety labels can be treated as culturally universal ground truth. The recurrent alternative is pluralistic evaluation: diverse annotator provenance, explicit geo-cultural modeling, and preservation rather than suppression of disagreement.
4. LLMs and the limits of default neutrality
Recent LLM audits consistently reject the idea that default model behavior is culturally neutral. “Cultural Bias and Cultural Alignment of LLMs” compares five GPT models against the Integrated Values Surveys using the Euclidean-distance measure
5
where smaller 6 indicates greater cultural alignment. All models cluster near countries such as Finland, Netherlands, Sweden, Norway, Denmark, Iceland, Australia, and New Zealand, and are furthest from countries such as Jordan, Libya, Tunisia, and Ghana. The dominant bias is toward self-expression values. Cultural prompting improves alignment for a majority of countries—71.0% for GPT-4o, 81.3% for GPT-4-turbo, and 77.6% for GPT-4—but it worsens or fails to help for 19–29% of countries and territories (Tao et al., 2023).
A broader multilingual study reaches a closely related conclusion. “LLMs and Cultural Values: the Impact of Prompt Language and Explicit Cultural Framing” probes 10 LLMs with 63 items in 11 languages and reports that the models remain anchored to a restricted set of cultural defaults: the Netherlands, Germany, the US, and Japan. The models produce “fairly neutral responses” on most topics, but this neutrality is selective rather than culture-free; on items involving environmental protection, homosexuality, abortion, and casual sex, they often adopt more progressive stances than human respondents. Prompt language and explicit cultural perspective both introduce variation, but explicit cultural framing improves alignment more than targeted prompt language, and combining both is “no more effective than cultural framing with an English prompt” (Bulté et al., 6 Nov 2025).
Work on regional models further shows that linguistic localization does not guarantee cultural neutrality or local alignment. “Fluent but Culturally Distant” evaluates five Indic and five global instruction-tuned models on values and practices and finds that Indic models do not align more closely with Indian cultural norms than global models on any of four tasks. The paper explicitly states that, across prompting strategies, the US lies closer to India than any Indic model, and on normalized GlobalOpinionQA all models show a statistically significant US tilt with 7; even the least US-tilting model, aryabhatta-8b, has mean nCAD of 8. The diagnosis is data-centric: translated English instruction data and limited culturally grounded corpora are too weak to offset Western-centric priors (Agarwal et al., 25 May 2025).
Mechanistic studies suggest that these failures are not merely surface effects. “Neuron-Level Analysis of Cultural Understanding in LLMs” identifies culture-general and culture-specific neurons accounting for less than 1% of all neurons, concentrated in shallow to middle MLP layers, and shows that suppressing them degrades cultural benchmarks by up to 30 points while general NLU remains largely unaffected. “Isolating Culture Neurons in Multilingual LLMs” similarly reports that, on average, 56.7% of culture-specific neurons are pure culture-specific and that these pure culture neurons explain about 76.3% of the effect of ablating all culture neurons (Yamamoto et al., 9 Oct 2025, Namazifard et al., 4 Aug 2025).
5. Translation, dialogue, media platforms, and multilingual generation
Cross-cultural communication tasks show that “neutrality” often appears as flattening rather than fairness. In dialogue generation, “Bridging Cultural Nuances in Dialogue Agents through Cultural Value Surveys” introduces cuDialog, built from OpenSubtitles2018 English subtitles with 13 culture labels, 5 genres, and 6 Hofstede cultural dimensions. The benchmark represents each example as 9, and culturally enhanced generation conditions decoding on predicted cultural dimensions. The reported result is that incorporating cultural value surveys improves alignment with references and cultural markers, with gains on BLEU, ROUGE-L, BERTScore, and Distinctiveness, especially when using multilingual pretrained models (Cao et al., 2024).
Translation studies make the neutrality problem explicit. “Towards Style Alignment in Cross-Cultural Translation” treats neutrality as a failure mode in which translations collapse toward a middling register rather than preserving culturally meaningful style. The paper formalizes ideal style preservation as
0
and measures corpus-level alignment by a correlation score 1. It reports that politeness variation shrinks in translation: native-text standard deviations are 0.23 for Spanish, 0.20 for Japanese, and 0.20 for Chinese, whereas translations into those languages have lower standard deviations of 0.17, 0.09, and 0.13. The proposed RASTA method yields up to 56% improvement in style alignment with less than 1.5% degradation in translation-quality metrics (Havaldar et al., 30 Jun 2025).
A domain-specific adaptation benchmark makes cultural neutrality an explicit evaluation axis. “Do LLMs Understand Wine Descriptors Across Cultures?” defines Cultural Neutrality as maintaining neutrality “to avoid provoking negative perceptions or reactions from the target culture consumers.” The canonical example is “earthy”: in English wine discourse it can be acceptable, but direct translation into Chinese as “土味” may sound dirty or unrefined, whereas “泥土气息” is treated as a more neutral rendering. The paper evaluates this criterion on a 1–7 scale alongside Cultural Proximity and Cultural Genuineness, showing that neutrality in translation is not identical to literalness or local vividness (Zou et al., 16 Sep 2025).
Digital platforms exhibit analogous limits. “Cultural Values and Cross-cultural Video Consumption on YouTube” uses daily top-50 video lists and shows that global access does not produce universal cultural convergence. In the main regressions over 58 countries, cultural values are stronger predictors of cross-cultural video consumption than GDP per capita, language centrality, or Internet penetration. For cultural closeness, the cultural model has adjusted 2 versus 0.244 for the non-cultural model, and for cultural betweenness 0.208 versus 0.061. The authors conclude that YouTube is globally accessible but not culturally neutral in its effects (Park et al., 2017).
Multilingual evaluation and multimodal generation reinforce the same pattern. “MCEval” spans 13 cultures and 13 languages, with 39,897 cultural awareness instances and 17,940 cultural bias instances, and shows that English-only success can hide major disadvantages elsewhere; one reported case is Swedish native-language performance dropping from 0.6 to 0.2, a 66.7% decrease, after CultureBank-based fine-tuning. In multilingual text-to-image generation, “Where Culture Fades” introduces CultureBench with 7,932 samples across 15 language/region groups and reports that current models often produce culturally neutral or English-biased outputs. The proposed zero-training and fine-tuned methods raise CultureVQA from strong baseline levels such as 25.13 for StableDiffusion 3.5 to 33.91 and 36.63, respectively, while preserving fidelity- and diversity-related metrics (Huang et al., 13 Jul 2025, Shi et al., 21 Nov 2025).
6. Critiques, controversies, and reformulations
A central contemporary critique is that proxy-based neutrality is too shallow. “‘Too much alignment; not enough culture’” argues that current approaches reduce culture to nationality, ethnicity, language, religion, race, values surveys, or lists of cultural facts. Against this, it proposes “thick outputs,” adapted from Clifford Geertz’s “thick description,” and states three necessary conditions for cultural alignment: sufficiently scoped cultural representation, capacity for nuanced outputs, and prompt anchoring in the cultural context implied by the interaction. The paper’s strongest conceptual claim is that cultural alignment is interaction-level rather than model-level: whether an output is culturally aligned can only be judged in a specific context (Orlowski et al., 30 Sep 2025).
A related position paper generalizes the critique from model behavior to benchmarking practice. “Culture is Everywhere: A Call for Intentionally Cultural Evaluation” argues that evaluation in language technology is never culturally neutral because culture is embedded in task selection, metrics, references, interaction settings, and interpretation. One concrete example is the claim that 28% of MMLU dataset requires culturally-sensitive knowledge to answer correctly. The proposed alternative is “intentionally cultural evaluation,” organized around what is evaluated, how it is evaluated, and the circumstances under which evaluation is conducted, with explicit attention to researcher positionality and participatory design (Oh et al., 1 Sep 2025).
Another controversy concerns whether neutrality can become erasure. In “Neutrality Bites: Gender Representation in AI-Generated Animal Stories,” Finkley, Li, and Walsh examine 23.8K stories and find that models avoid gendering the protagonist in 19% of stories, use gender-neutral language in 38.2%, feature masculine characters in 40.6%, and feminine characters in only 2.2%. Their conclusion is that neutrality can suppress identity without producing equitable representation: the output looks less marked, but the imagined world remains heavily masculinized (Finkley et al., 6 Jun 2026).
These critiques do not converge on a single replacement for neutrality. Some advocate pluralistic safety alignment, some contextual prompting, some explicit cultural conditioning, some ethnographic evaluation, and some mechanism-level editing. What they share is a rejection of the claim that neutrality can be achieved simply by stripping culture away. The dominant reformulation is narrower and more conditional: neutrality, if retained as a useful ideal, denotes non-dominance, explicit scope, and context-sensitive handling of plural values rather than a universal culture-free standard.