Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cultural Embeddings (CuE)

Updated 10 July 2026
  • Cultural Embeddings (CuE) are approaches that represent culture-sensitive meaning as relational geometry in continuous embedding spaces.
  • CuE methodologies employ techniques like hyperbolic embeddings, semantic directions, and sparse LLM features to analyze cultural dynamics.
  • They bridge computational text analysis with sociological theory, enabling measurable insights into symbolic boundaries and semantic change.

Searching arXiv for the cited CuE-related papers to ground the article in current arXiv records. Cultural Embeddings (CuE) designate a family of approaches that use embedding spaces to represent culture-sensitive meaning as relational geometry rather than as isolated traits or raw keyword counts. In the foundational sociological formulation, word embeddings are treated as “cultural embeddings”: standardized, continuous “meaning space[s]” in which distance is meaningful, semantic change can be tracked, and symbolic boundaries can be measured (Stoltz et al., 2020). Subsequent work extends this general idea in several directions: hyperbolic embeddings for jointly mapping social and semantic hierarchy, contrastively trained spaces for culture-specific common ground, multimodal embeddings derived from sketches or text-to-image systems, and sparse feature-level representations inside LLMs used for diagnosis and steering (Wu et al., 2018). The common thread is that cultural structure is modeled as geometry—graded, relational, and often historically or interactively variable—while interpretation remains essential rather than replaceable (Stoltz et al., 2020).

1. Relational meaning and the original CuE formulation

The original CuE framework is articulated most explicitly in “Cultural Cartography with Word Embeddings” (Stoltz et al., 2020). Its starting point is the claim that frequency-based text analysis glosses over the relationality of word meanings, whereas embeddings place words in a continuous semantic space where words that appear in similar contexts are located near one another. This is aligned with the distributional hypothesis associated with Zellig Harris and Firth’s formulation that “you shall know a word by the company it keeps,” and it is tied to sociological work by Mohr, Breiger, White, Mische, Zelizer, and others (Stoltz et al., 2020).

In this formulation, CuE is not a substitute for interpretation. It is a way to map the relational structure of meaning in texts with more sensitivity than simple word counts. The basic claim is that cultural meaning is often graded and relational, not merely present or absent. Counting how often a term appears does not capture whether it is associated with other concepts, social classifications, or evaluative poles, whereas embeddings represent those shifts as distances in a low-dimensional vector space produced by dimension reduction of a document-term matrix or a term-co-occurrence matrix (Stoltz et al., 2020).

A closely related contribution appears in “The Geometry of Culture: Analyzing Meaning through Word Embeddings” (Kozlowski et al., 2018). That paper recasts word embeddings as a tool for cultural sociology by treating word-difference vectors such as manwoman\text{man}-\text{woman}, richpoor\text{rich}-\text{poor}, blackwhite\text{black}-\text{white}, and liberalconservative\text{liberal}-\text{conservative} as dimensions of cultural meaning. A cultural dimension is defined as the normalized mean difference over a set of antonym pairs, and a word’s cultural association is measured by its projection onto that dimension, which is equivalent to cosine similarity when vectors are normalized (Kozlowski et al., 2018).

This line of work positions CuE as a methodological bridge between computational text analysis and relational theories of meaning. The emphasis is not on one universal cultural axis, but on a geometry with multiple partially independent dimensions that can align, drift apart, or vary across corpora and historical periods. “High-dimensional theorizing,” in this sense, is not merely computational convenience; it is presented as a way to preserve intersectionality, heterogeneity, and historical change in the analysis of culture (Kozlowski et al., 2018).

2. Geometries, coordinates, and comparative spaces

The most common CuE geometry in the literature is the vector space of word embeddings, where words receive vectors vw\mathbf{v}_w and similarity is commonly measured with cosine similarity on a scale from 1-1 to $1$ (Stoltz et al., 2020). Within this setting, two geometric objects are especially important. A semantic direction is a vector extracted by subtracting one set of terms from another, such as “poverty” from “affluence,” while a semantic region can be a centroid or cluster of related terms (Stoltz et al., 2020). These constructions support analyses of symbolic boundaries, social marking, and conceptual engagement.

A different geometry appears in “Social Connection Induces Cultural Contraction” (Wu et al., 2018), which embeds both social networks and semantic or topic networks into a two-dimensional hyperbolic space, specifically a Poincaré disk. The motivation is structural: real social and semantic systems are hierarchical, sparse, clustered, and full of bridging ties, whereas Euclidean space is described as ill-suited because it enforces transitivity too strongly and poorly represents tree-like or scale-free structures. In the Poincaré disk, the radial coordinate rr represents hierarchy or centrality, and the angular coordinate θ\theta represents similarity, divergence, or diversity (Wu et al., 2018).

The hyperbolic formulation is important because it makes social actors and cultural items geometrically comparable. Universities are embedded in collaboration space, PACS physics topics are embedded in topic co-occurrence space, and universities are then projected into cultural space by summarizing the PACS codes in their publications. The paper summarizes a university’s cultural diversity by the entropy of its PACS-code angle distribution and its cultural hierarchy by its radius, thereby turning hierarchy tightening and angular spread into directly measurable properties (Wu et al., 2018).

More recent work broadens CuE beyond text-only geometry. “Billions of Sketches Reveal Hidden Cultural Variation in Human Concepts” (Pera et al., 8 Jul 2026) constructs sketch-based embeddings from 2.6 billion QuickDraw sketches of 344 concepts from 236 countries and territories. Each reconstructed sketch image is embedded with DINOv2 into a 384-dimensional feature vector, reduced to 40 dimensions by PCA and then to 2 dimensions by UMAP for visualization and clustering. Within-concept clusters are identified with DBSCAN, with ϵ\epsilon chosen by DBCV optimization, and clusterability is defined as the proportion of non-noise sketches (Pera et al., 8 Jul 2026). In this setting, a concept is not a single point but a set of visual exemplars.

This suggests that CuE is not tied to one geometry or one modality. Euclidean word spaces, hyperbolic disks, sketch-derived concept spaces, and feature spaces inside LLMs all instantiate the same broader move: representing culture-sensitive structure as relations among points, directions, regions, or clusters rather than as isolated labels.

A central distinction in CuE methodology is between two “navigation” modes introduced in the sociological literature (Stoltz et al., 2020). The first is holding terms constant while the embedding space moves. This variable embedding space strategy splits the corpus by a covariate—usually time, but also author, community, or organization—and trains a separate embedding on each subset. The same word or concept is then compared across spaces. The second is holding the embedding space constant while documents or authors move relative to it. This fixed embedding space strategy trains a single embedding on a corpus and represents each document or author as a cloud of points or a distribution within that common space (Stoltz et al., 2020).

The empirical illustration in U.S. immigration discourse uses both strategies. For variable embedding spaces, decade-specific embeddings from the Corpus of Historical American English are used from 1880 to 2000 to examine how “immigration” changes in relation to “job,” “crime,” “family,” and “school,” as well as to race, social class, and morality dimensions adapted from Kozlowski et al. (2019). For fixed embedding spaces, the “All the News” corpus of about 204,135 U.S. news articles from 18 outlets, mostly from 2013 to early 2018, is compared to 986 press releases from four immigration-focused advocacy organizations, using Word Mover’s Distance (WMD) and Concept Mover’s Distance (CMD) (Stoltz et al., 2020).

The associated analytic operations are geometrically explicit. WMD compares two documents by the minimal “transport cost” required to move one document’s word mass onto another, using embedding distances as costs. CMD compares documents to semantic directions or pseudo-documents representing concepts, thereby measuring the cost of moving a document’s word distribution toward a concept such as “immigration + school” or “immigration + family” (Stoltz et al., 2020). These are transportation-style measures rather than count-based indicators.

“The Geometry of Culture” adds another important operation: projection-based connotation measurement and stability testing (Kozlowski et al., 2018). After constructing gender, class, and race dimensions by averaging antonym differences, the paper measures a word’s cultural position by projecting onto those directions. It then estimates uncertainty through bootstrap and subsampling procedures: repeatedly resampling the corpus, retraining embeddings, recomputing projections and distances, and forming empirical confidence intervals, often reported as 90% bootstrap confidence intervals (Kozlowski et al., 2018).

A related but distinct workflow appears in “Machine learning as a model for cultural learning” (Arseniev-Koehler et al., 2020), where embeddings trained on 103,581 New York Times articles from January 1, 1980 to July 15, 2016 are used to extract cultural schemata about body weight. The selected model uses 500-dimensional vectors, context window richpoor\text{rich}-\text{poor}0, CBOW, and negative sampling with 5 negative samples, with a reported Spearman correlation of .61 on WordSim353. Higher-order schemata are derived as semantic directions for gender, morality, health, and socioeconomic status, and then used to show that obesity-related terms align with feminine, immoral, unhealthy, and low-SES poles (Arseniev-Koehler et al., 2020).

Across these workflows, CuE provides not one algorithm but a repertoire of operations: train spaces, align spaces, define semantic directions, compute cosine projections, measure mover distances, resample for stability, and compare documents, concepts, or institutions inside a common geometry.

4. Substantive findings in cultural sociology and social theory

The sociological CuE literature uses embeddings to study semantic change, social marking, media fields, echo chambers, and cultural diffusion (Stoltz et al., 2020). In the historical immigration analysis, “immigration” becomes increasingly associated with “crime” over the 19th and 20th centuries, while its association with “job,” “school,” and “family” changes more modestly. When “immigrant” and “citizen” are compared to race, class, and morality dimensions, citizens consistently cluster closer to the white, high-class, and good poles, while immigrants tend closer to the black, low-class, and bad poles. This is interpreted as either the stability of unmarked normative categories or a durable symbolic boundary in U.S. discourse, specifically the persistent association of “immigrant-as-nonwhite” versus “citizen-as-white” (Stoltz et al., 2020).

At the document level, left-leaning news organizations become more similar to both left- and right-wing immigration advocacy press releases over time, especially in 2016–2017, while right-leaning outlets are consistently more similar to right-wing advocacy texts when discussing immigration. The paper reads this as consistent with echo chambers, asymmetric polarization, and heteronomy in the journalistic field, with attention to DACA, DAPA, Trump’s immigration rhetoric, and the 2016 election. CMD further shows that media discourse on “immigration” is closely linked to compound concepts such as “immigration + school” and “immigration + family,” with peaks around DACA and DAPA debates (Stoltz et al., 2020).

“The Geometry of Culture” provides complementary historical and comparative findings (Kozlowski et al., 2018). Using decade-specific Google Ngrams embeddings for U.S. books from 1900–1999, it shows that “nurse” becomes gradually less feminine over time, “engineer” becomes gradually less masculine, and “journalist” shifts from masculine toward feminine, especially later in the century. Along the class dimension, “journalist” becomes markedly more upper-class, “nurse” rises somewhat in class, and “engineer” stays relatively middle-class. The paper also tracks the angle between gender and class vectors, arguing that these dimensions move toward orthogonality over the 20th century (Kozlowski et al., 2018).

The same paper compares U.S. and British embeddings from roughly 1890–1910 and finds that some class and gender neighborhoods are shared, while the fine structure differs. In Britain, “bare-covered” and “employed-unemployed” are more closely aligned with class than in the U.S. in that period; British texts also more strongly feminize colonized places such as Africa, Asia, and India (Kozlowski et al., 2018).

“The dynamics of cultural systems” generalizes the relational idea beyond word vectors into a systems-level account of culture (Jansson, 1 Jan 2026). There, CuE is not defined as one particular embedding model but as a way of thinking about culture as dynamically embedded in cognitive, social, and material ecologies. Cultural traits sit inside mental schemas, institutions, infrastructures, and artefacts; coherence-seeking information processing, filtering, path dependence, attractor states, epistemic niches, and feedback loops structure cultural evolution. In that view, LLMs and recommender systems are products of cultural embeddings that now act back on cultural systems by filtering and recombining information (Jansson, 1 Jan 2026).

A common misconception is that CuE is only a visualization technique. The literature does not support that reduction. In the sociological papers, embeddings are used to operationalize symbolic boundaries and connotations; in systems work, they are tied to path dependence, coherence, and filtering; and in later machine-learning work they become instruments for routing, mining, and white-box steering.

5. Social connection, contraction, and cross-cultural common ground

One major extension of CuE concerns the relation between social structure and cultural geometry. In the physics case studied with Poincaré embeddings, 359,395 APS papers, 372,495 authors, 28,754 institutions, and 5,819 PACS codes are embedded year by year from 2002–2011 (Wu et al., 2018). The institutional network reveals a collaboration hierarchy, the PACS network reveals a topical hierarchy, and the angular structure shows strong topical clustering. Across universities, social distance and cultural distance are strongly positively correlated over time; the correlation between the speed of social convergence or divergence and the speed of cultural convergence or divergence is reported as richpoor\text{rich}-\text{poor}1, richpoor\text{rich}-\text{poor}2 (Wu et al., 2018).

The paper’s central claim is that denser communication channels increase shared information and therefore reduce the overall entropy of the system: the cultural space available to the group contracts. Greater external collaboration predicts future convergence toward globally popular topics, and higher clustering or density is associated with lower entropy of PACS codes. The paper therefore argues that ties can facilitate bricolage and new combinations locally while accelerating the extinction of alternative forms globally (Wu et al., 2018).

A different aspect of cross-cultural structure appears in “Communicate to Play” (White et al., 2024), where culturally informed embeddings model culture-specific common ground in Codenames Duet. The representation begins from GloVe vectors and learns a linear transformation

richpoor\text{rich}-\text{poor}3

Training uses a contrastive or cross-entropy objective that encourages high similarity between the embedded clue and human-selected words while lowering similarity to distractors, with scores based on cosine similarity scaled by richpoor\text{rich}-\text{poor}4 (White et al., 2024). The resulting space is intended to encode clue–target associations that humans from a given subset actually use, including culture-specific regularities.

These embeddings are then inserted into RSA+C3, a pragmatic reasoning system with multiple literal listener models richpoor\text{rich}-\text{poor}5, one for each cultural group. The clue giver updates beliefs about the interlocutor’s cultural context using a smoothed memory term and then chooses clues with a speaker objective proportional to

richpoor\text{rich}-\text{poor}6

On the Cultural Codes dataset of 794 Codenames Duet games and 153 players, trained embeddings outperform untrained GloVe and often few-shot Llama2 prompting on human gameplay tasks; for guess accuracy on all data, GloVe baseline guess accuracy is 43.16 and trained embeddings reach 60.50, a 40.18% improvement. Across splits such as education and country, models trained on the same cultural subset perform better than those trained on a different subset, and RSA+C3 achieves the highest win rate in the reported cross-cultural interactive settings (White et al., 2024).

This suggests that CuE can operationalize two related but distinct propositions. First, social integration can compress the diversity of cultural possibilities. Second, pragmatic failure can arise because speakers and listeners inhabit different local semantic neighborhoods, and adaptation requires modeling that difference rather than assuming a universal common ground.

6. CuE in LLM alignment, multimodal representation, and cultural localization

Recent work applies CuE to LLM alignment and multimodal generation in ways that differ from the original sociological formulation but preserve the emphasis on latent geometric structure. In “Whispers of Many Shores” (Feng et al., 30 May 2025), cultural expertise is represented without full fine-tuning of the base model by a modular soft-prompt or embedding-based routing framework. User profile, questionnaire answers, and option variables are encoded using mxbai-embed-large; a “Sentopic Agent” flags culturally sensitive content with an LLM-as-judge prompt and the paper reports 95% accuracy for this detection step; topic extraction, planning, top-richpoor\text{rich}-\text{poor}7 cultural expert routing, and a Composer Agent then produce a final response (Feng et al., 30 May 2025).

The routing formalism fuses a topic centroid and a user embedding into

richpoor\text{rich}-\text{poor}8

computes similarity to expert embeddings richpoor\text{rich}-\text{poor}9 by negative blackwhite\text{black}-\text{white}0 distance, applies Top-blackwhite\text{black}-\text{white}1 selection, uses a clustering-based fallback when the best score is below threshold blackwhite\text{black}-\text{white}2, and assigns softmax-normalized expert weights. In experiments on 100 simulated user profiles spanning 20 countries and using IBM Granite 3.3, the reported Cultural Alignment Score improves from 0.208 to 0.820, diversity entropy rises from 0.443 to 1.659, and latency increases from 5.843 s to 44.931 s (Feng et al., 30 May 2025).

“Steering LLMs for Culturally Localized Generation” introduces a white-box CuE based on sparse autoencoders (SAEs) (Khanuja et al., 24 Mar 2026). Here, CuE is a country-level sparse vector built from SAE features selected by mutual information with country labels over 4,334 assertions derived from CANDLE and additional culturally salient assertions. For each input blackwhite\text{black}-\text{white}3, max-pooled activations

blackwhite\text{black}-\text{white}4

are restricted to a selected feature set blackwhite\text{black}-\text{white}5, and country prototypes are formed by averaging over country-specific assertion sets. Centered cosine similarity between a generation and a country prototype is then used to diagnose implicit cultural bias, while a contrastive difference between a target country prototype and the mean of all others is decoded through the SAE decoder into a residual-stream steering vector (Khanuja et al., 24 Mar 2026).

That paper reports that under implicit prompting, Gemma-2-9B defaults heavily toward Anglo cultures: 33.1% of outputs align most closely with the U.S. and 26.9% with the U.K. The bias concentration index falls from 0.333 under implicit prompting to 0.051 under blackwhite\text{black}-\text{white}6 and 0.045 under blackwhite\text{black}-\text{white}7. For Gemma-2-9B 16K, blackwhite\text{black}-\text{white}8 versus explicit prompting yields 48% wins and 24% losses on cultural faithfulness and 53% wins and 17% losses on rarity; blackwhite\text{black}-\text{white}9 versus explicit prompting yields 55% wins on faithfulness and 59% wins on rarity (Khanuja et al., 24 Mar 2026).

Two additional lines of work extend CuE toward data curation and supervision. “C-Mining” treats cultural specificity as a measurable property of multilingual embedding spaces, exploiting cross-lingual misalignment, local density, and linguistic dominance to extract Culture Points from multilingual Wikipedia without human or LLM supervision. Its criterion

liberalconservative\text{liberal}-\text{conservative}0

uses liberalconservative\text{liberal}-\text{conservative}1 and liberalconservative\text{liberal}-\text{conservative}2 for cluster selection; mined seeds are then used to synthesize 50,000 instruction-response pairs, producing a +6.03 gain on CulturalBench-Hard for Qwen2.5-7B and more than 150-fold preparation-cost reduction relative to manual curation (Zeng et al., 17 Apr 2026). “From National Curricula to Cultural Awareness” does not define an explicit CuE architecture, but it provides culture-grounded supervision through CuCu and KCaQA: 34,128 open-ended QA pairs derived from 158 Korean social studies learning outcomes, with 2,844 query variants and three response levels per query (Yoo et al., 8 Jan 2026).

Multimodal CuE also appears in studies of concept representation and generation. Sketch-derived concept embeddings preserve exemplar structure and align 45% more closely with established cultural distances than text-based measures, with a macro average rank correlation of 0.098 between visual and linguistic similarity rankings (Pera et al., 8 Jul 2026). In text-to-image models, culture is probed through a three-tier ontology of dimensions, domains, and concepts and through multilingual prompt templates such as “a photo of <concept>,” translated prompts, nationality-tagged prompts, and even gibberish strings in target alphabets. Intrinsic evaluations in CLIP space and extrinsic evaluations with BLIP2 indicate that text-to-image models do encode cultural information and that culture can sometimes be unlocked by explicit nation prompts or even script-level cues (Ventura et al., 2023).

7. Limits, controversies, and open directions

The CuE literature is explicit that embeddings do not “understand” meaning in any human sense; they reflect patterns in text, images, sketches, or model activations, and interpretation remains essential (Stoltz et al., 2020). Results depend on the representativeness and quality of the training corpus, pretrained embeddings may omit marginalized voices or communities, and independently trained variable embedding spaces require careful alignment because they are not directly comparable (Stoltz et al., 2020). In the obesity study, the use of a single newspaper, a specific discourse domain, and a single vector per word limits what can be claimed about broader or context-sensitive meaning (Arseniev-Koehler et al., 2020).

Several controversies follow from these limitations. One concerns reductionism: binary oppositions and semantic directions are only one type of relation. Hierarchy, entailment, and part-whole structure may require different methods or even different geometry, such as hyperbolic embeddings (Stoltz et al., 2020). Another concerns essentialization: some LLM-localization work uses country as a proxy for culture, which is explicitly described as pragmatic but flattening, since it ignores within-country diversity and risks essentializing identity (Khanuja et al., 24 Mar 2026). A third concerns the difference between missing knowledge and poor elicitation. White-box steering results suggest that at least some long-tail cultural knowledge is already latent inside models and can be surfaced by better elicitation, though the strength of this effect varies across cultures (Khanuja et al., 24 Mar 2026).

There is also a modality controversy. Linguistic analyses often treat conceptual similarity across languages as evidence of universality, but sketch-based work argues that words are “lossy” because they compress rich experiential variation into shared labels. The divergence between sketch and word geometries, and the stronger alignment of image-based country networks with cultural-distance measures, suggests that conclusions about universality may depend on the modality through which concepts are measured (Pera et al., 8 Jul 2026).

The current trajectory of CuE research points toward integration rather than replacement. A plausible implication is that future work will combine variable and fixed embedding strategies, contextual or multimodal representations, culture-grounded supervision, and mechanistic interventions. That direction is already visible across the literature: sociological cartography of meaning, hyperbolic modeling of hierarchy and diversity, cross-cultural pragmatic reasoning, multilingual seed mining, curriculum-grounded QA generation, and white-box steering all treat culture as structured and measurable in latent space, while differing on what is embedded, which geometry is appropriate, and how much of cultural variation can be operationalized without collapsing it into stereotype.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Cultural Embeddings (CuE).