Papers
Topics
Authors
Recent
Search
2000 character limit reached

GatheringSense: AI-Generated Imagery and Embodied Experiences for Understanding Literati Gatherings

Published 13 Feb 2026 in cs.HC | (2602.12565v1)

Abstract: Chinese literati gatherings (Wenren Yaji), as a situated form of Chinese traditional culture, remain underexplored in depth. Although generative AI supports powerful multimodal generation, current cultural applications largely emphasize aesthetic reproduction and struggle to convey the deeper meanings of cultural rituals and social frameworks. Based on embodied cognition, we propose an AI-driven dual-path framework for cultural understanding, which we instantiate through GatheringSense, a literati-gathering experience. We conduct a mixed-methods study (N=48) to compare how AI-generated multimodal content and embodied participation complement each other in supporting the understanding of literati gatherings and fostering cultural resonance. Our results show that AI-generated content effectively improves the readability of cultural symbols and initial emotional attraction, yet limitations in physical coherence and micro-level credibility may affect users' satisfaction. In contrast, embodied experience significantly deepens participants' understanding of ritual rules and social roles, and increases their psychological closeness and presence. Based on these findings, we offer empirical evidence and five transferable design implications for generative experience in cultural heritage.

Summary

  • The paper introduces GatheringSense, a dual-path framework that combines AI-generated multimodal scenes with physically enacted rituals to support cultural understanding through seeing, perceiving, and resonating.
  • In a mixed-methods study of 48 participants, embodied experience significantly increased cultural resonance from 4.73 to 5.54 and psychological closeness from 3.00 to 4.66, while presence strongly correlated with resonance (ρ=.66, p<.001).
  • The findings show that ink-wash and fine-brush styles improve semantic clarity, while personalization, ritual participation, social interaction, and exploration matter more to user satisfaction than extensive knowledge guidance or real-time AI feedback.

Overview and motivation

This paper addresses a persistent gap in digital cultural heritage: systems can render cultural symbols visible, but rarely convey the procedural rules, role structures, and social atmospheres that constitute complex cultural practices. The authors ground their argument in Geertz's notion of "thick description" and in embodied cognition theory, arguing that practices such as Chinese literati gatherings (Wenren Yaji)—which interweave poetry, calligraphy, painting, music, and tea ceremony within ritualized turn-taking and spatial choreography—cannot be understood through observation alone. Their response is an AI-driven dual-path framework combining (1) an AI symbolic path, in which AI-generated multimodal content makes scenes, activities, and roles recognizable, and (2) an embodied experience path, in which participants physically enact the gathering's behavioral logic. The two paths are linked by a three-stage cognitive relay: "Seeing – Perceiving – Resonating." The framework is instantiated as GatheringSense and evaluated in a mixed-methods study (N=48N=48) with three research questions covering initial understanding from AI content (RQ1), embodied contributions to immersion and social presence (RQ2), and how the paths jointly foster cultural resonance (RQ3).

The paper's positioning against prior work is well supported. Existing generative-AI cultural applications emphasize aesthetic reproduction and lack cultural semantic alignment (2602.12565), while immersive heritage systems focus on artifact reconstruction rather than structured social activity patterns. GatheringSense is distinctive in explicitly designing for social presence, ritual participation, and role transitions, and in quantifying these with validated instruments.

Experience design

The study corpus draws on five canonical texts spanning the Wei–Jin through Yuan periods, including the Preface to the Orchid Pavilion Gathering and the Record of the Elegant Gathering in the Western Garden. Four core activities—pitch-pot (touhu), Go, calligraphy, and poetry singing—were selected for historical frequency, symbolic value, and feasibility of physical enactment.

The AI visualization pipeline used GPT-5 for semantic extraction, then compared four text-to-image models (gpt-image-1, Dou Bao AI, Stable Diffusion 3.5, Flux Pro 1.1 Ultra), finding Dou Bao AI strongest for ink-wash and fine-brush styles and gpt-image-1 for oil and cartoon styles. Sixteen key frames (four activities × four styles) were interpolated into 30–40 s videos via Jimeng 2.0, with AI-generated guqin/xiao soundscapes from Suno. Notably, the authors report that generation models struggled with pitch-pot—a culturally unfamiliar, physically dynamic activity—and its video was subsequently poorly received, an honest early signal of AIGC limits on procedural content.

The embodied path was staged in a laboratory configured as a gathering space with a curved screen, circular seating, structured prop-passing paths, ambient music, and role rotation between performer and observer. Substitutes preserved symbolic fidelity while ensuring safety (e.g., Gomoku replaced Go for session-duration reasons).

Method

Participants were 48 people divided into four groups: Chinese adults in forward order (G1, N=23N=23; AI then embodied), Chinese adults in reverse order (G2, N=15N=15; embodied then AI), cross-cultural adults (G3, N=4N=4), and children aged 7–10 (G4, N=6N=6). Only G1+G2 entered inferential analyses; G3/G4 were exploratory. The sample was deliberately "technologically familiar but culturally unfamiliar" (AIGC experience M=4.57M=4.57, literati familiarity M=3.24M=3.24 on 7-point scales), which suits testing the framework under low prior knowledge.

Measures included a Film-IEQ proxy for interpretability of images vs. videos, a custom Cultural Resonance (CR) scale administered at three stages, selected ITC-SOPI items for embodied presence, IOS for psychological closeness to the cultural concept, SAM at four time points, KANO questionnaires over ten candidate features, association-keyword elicitation, and semi-structured interviews analyzed via reflexive thematic analysis. Quantitative analysis used linear mixed models with planned contrasts and Tukey corrections.

Quantitative results

Media format and style. Static images outperformed videos on the Film-IEQ proxy (M=5.28M=5.28, SD=1.01SD=1.01 vs. M=5.04M=5.04, N=23N=230; N=23N=231, N=23N=232), indicating that key frames clarify "who is doing what" better than short clips, while video contributes mainly dynamic and affective information. Style rankings strongly favored ink-wash (N=23N=233) and fine-brush (N=23N=234) over oil (N=23N=235) and cartoon (N=23N=236). The authors attribute this to style–semantic alignment: culturally congruent brushwork lexicons reduce cognitive load, whereas cross-cultural or modern abstractions weaken semantic fit. This yields a concrete design rule—use ink-wash/fine-brush as primary narrative carriers and reserve oil/cartoon for contrast.

Resonance trajectory. CR was already mid-high after AI media (CR_img N=23N=237; CR_vid N=23N=238; difference n.s.), but rose significantly to N=23N=239 post-experience (N=15N=150, N=15N=151). Embodied presence correlated strongly with post-study resonance (N=15N=152, N=15N=153), while prior familiarity with literati gatherings did not (N=15N=154, N=15N=155). The implication is notable: participants initially unfamiliar with the practice could reach high cultural resonance given sufficient embodied presence, supporting the claim that embodiment—not prior knowledge—is the operative lever.

Psychological closeness and order effects. IOS increased from N=15N=156 to N=15N=157 overall (N=15N=158, N=15N=159). Critically, the gain was larger when embodiment preceded AI media (G2: N=4N=40) than the reverse (G1: N=4N=41; N=4N=42, N=4N=43), suggesting embodied experience has a stronger initiating effect on reducing psychological distance, with subsequent AI content consolidating structure. However, the authors caution that per-group sample sizes were modest, so order-effect conclusions should be interpreted cautiously—an appropriate concession given the small G2.

Affect. SAM remained stable across phases (valence 4.16–4.68, arousal 4.89–5.24), which the authors frame as intentional: the system targets a stable emotional background rather than arousal peaks, with change occurring in structural understanding and identification instead.

KANO analysis. Seven of ten features (F2–F8) were classified Attractive; none were Must-be or One-dimensional. Personalization had the highest N=4N=44 (78.57%), followed by ritual participation (78.05%) and environmental immersion/exploration freedom (both 76.19%). Social interaction and exploration freedom carried the largest dissatisfaction coefficients (N=4N=45). Content breadth, knowledge guidance, and AI real-time feedback were Indifferent (e.g., F9 N=4N=46). The strong implication is that users do not want didactic knowledge delivery; satisfaction hinges on agency, ritual, and social structure—"being one of the gathering"—rather than on visible AI capability or explanatory completeness.

Qualitative results

Association data (3,214 word occurrences across G1+G2) showed that both groups first constructed static scenes around environment (~30% of core words) and material media (~27%), but shifted toward specific cultural identities ("literati," "scholars"), nuanced affective vocabulary ("refined," "secluded"), and activity-specific action verbs ("composing," "chanting," "wielding")—independent linguistic evidence for the seeing-to-resonating relay.

Interviews reinforced a clear division of labor. AI content served as an entry point and "problem space": participants valued ink-wash/fine-brush for elegance and contextual fit but were highly sensitive to physical incoherence—clipping, distorted hand positions, motions violating real-world logic—with one participant summarizing, "It gives me a rough feeling; I treat the details as AI bugs." After embodied experience, tolerance improved: one participant reported that having performed pitch-pot, they could "guess what it wants to express" even when the video was inaccurate. The embodied path built presence through ritual and social interaction—Go etiquette read as "a kind of Li," pitch-pot as a "social catalyst"—amplified by setting and music. Participants also acknowledged limits: some stayed in a "playful, appreciative attitude" without fully entering the historical context, indicating that fun alone does not produce deep understanding without scripting and guidance.

Exploratory findings suggest transferability with reweighting. Cross-cultural participants treated text/AI as a "necessary but effortful manual" and activities as the experiential core; children dismissed AI videos as "funny background" (with sharp critiques such as "AI is not ready yet, it's very unsafe") and responded to embodiment as a "universal language." These are small samples (N=4N=47 and N=4N=48) with no inferential statistics, so they should be read as directional only.

Discussion and design implications

The discussion frames AI-generated visual content as "a bridge full of gaps": effective for scene-level semantic scaffolding, but prone to uncanny-valley dissonance in dynamic scenarios due to failures in physical credibility, temporal coherence, and emotional expression. The style findings motivate a positioning of AIGC as a "preliminary framework" rather than final representation, favoring stylized modes that disclose uncertainty over hyperrealism. The central empirical contribution is quantified evidence for the embodied path: the +1.66 IOS gain and the presence–resonance correlation align with situated-action accounts of cognition and distinguish this work from single-artifact, single-user heritage systems.

From the KANO structure, the authors derive five design probes: (A) physical credibility of visual symbols; (B) visualization of social frameworks and turn-taking mechanisms; (C) multisensory atmosphere; (D) ritualization of behavioral flow; and (E) AI content explainability and style management. Probe E is particularly actionable—it converts AI limitations into an acceptable stylistic register rather than a defect, directly addressing the negative reactions observed in the video condition.

Limitations and open questions

The paper is candid about several constraints. The cross-cultural and child groups are too small for inference; the order counterbalancing involved modest per-group sizes, limiting power to detect subtle order effects despite the significant IOS order result. The embodied setting was an indoor prototype rather than a historically reconstructed context, leaving open whether effects hold in gamified VR or situated deployments in museums and schools. The AI content reflects general-purpose models' current weaknesses in physical plausibility; targeted cultural-activity datasets or customized pipelines remain untested. An additional open question concerns durability: all measures were collected within a single ~60-minute session, so whether gains in closeness and resonance persist beyond the session is not established.

Conclusion

GatheringSense provides empirical support for treating AI-generated multimodal content and embodied participation as complementary rather than substitutive paths in cultural understanding. AI media reliably raise symbol readability and establish a mid-high baseline of cultural resonance; embodied, socially structured interaction produces the substantial gains in presence, psychological closeness, and resonance, with an order advantage when embodiment precedes AI viewing. The KANO results redirect design attention from didactic knowledge delivery toward personalization, ritual participation, and social exploration. The main caveats—small supplementary groups, limited statistical power for order effects, prototype-level embodiment, and current AIGC physical-coherence deficits—define the immediate agenda for validating and extending the dual-path framework in more diverse populations and deployed cultural settings.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.