---
title: 'HAIVMet: Human–AI Visual Metaphors'
url: https://www.emergentmind.com/topics/haivmet-human-ai-visual-metaphors
type: topic
---

# HAIVMet: Human–AI Visual Metaphors

Human–AI Visual Metaphors (HAIVMet) designate a computational paradigm and family of systems that leverage the complementary interpretive powers of humans and artificial intelligence to generate, analyze, transfer, or explain nonliteral meaning in visual modalities. Rooted in cognitive metaphor theory and extended through advances in generative modeling, schema reasoning, and collaborative design interfaces, HAIVMet reframes both the construction and interpretation of visual metaphors, enabling richer, more actionable connections between abstract concepts and concrete imagery. Contemporary HAIVMet systems integrate large language models, vision–language models, and interactive agents with human-in-the-loop workflows across disparate domains, including dialogue visualization, creative image synthesis, video captioning, scientific visualization, and affective design [2508.07520, 2602.01335, 2305.14724, 2504.10101, 2508.18569, 2403.00632, 2507.22051].

## 1. Theoretical Foundations and Metaphorical Mapping

HAIVMet instantiates foundational principles from cognitive linguistics and visual rhetoric by operationalizing core constructs such as cross-domain mapping, conceptual blending, and schema invariance. Systems in this class encode the abstract structure of a metaphor—typically as a mapping between source and target domains or as a triplet (primary concept, relation, secondary concept)—and instantiate this mapping in visual modalities via algorithmic or agentic means [2212.09898, 2502.16062, 2303.15445]. 

For example, in the schema-driven VMT (Visual Metaphor Transfer), a reference image is abstracted into a universal schema grammar encoding entities, generic-space relations (e.g., ontological, functional), aesthetic attributes, and violation points. The system then recomposes the abstract logic onto a distinct subject, ensuring the invariant relational mapping persists [2602.01335]. Analogously, in conversational visualization (Conversational DNA), linguistic features such as contribution volume, emotional valence, and semantic alignment are recast as intertwined helical strands, making the architecture of dialogue “seeable” as a living biological metaphor [2508.07520].

## 2. Computational and Algorithmic Frameworks

HAIVMet workflows blend symbolic reasoning, neural generation, and joint human–AI curation. Central methodologies include:

- **Schema Extraction and Grammar Encoding**: Systems abstract high-level logic from visual or textual input using parsers, large language models, and CLIP-embedding pipelines to identify component roles, attributes, and analogical invariants [2602.01335, 2502.16062, 2212.09898].
- **Agent-based Reasoning**: The VMT framework allocates specialized agents for perception (schema extraction), transfer (invariant preservation and adaptation), generation (prompt synthesis), and hierarchical diagnostics. Interaction between these agents enables closed-loop correction and robust metaphor reification [2602.01335].
- **Prompt Decomposition**: Structured prompting and S–T–M (Source–Target–Meaning) frameworks force explicit delineation of metaphorical structure, which is then injected into prompt engineering for text-to-image models [2508.18569, 2305.14724]. LaTeX-formatted reward functions (\(\mathcal{R}(I, P, D)\)) aggregate decomposition quality, CLIPScore, and meaning-alignment metrics to guide prompt refinement.
- **Human–AI Collaboration**: HAIVMet systems typically embed human review/validation steps. These include chain-of-thought (CoT) prompting for visual elaboration followed by expert or crowd-based filtering, as in the HAIVMet dataset [2305.14724], collaborative metaphor suggestion and editing in affective interfaces [2403.00632], or interactive annotation and feedback in understanding pipelines [2303.15445, 2212.09898].
- **Joint Multimodal Reasoning**: Architectures span vision–language transformers for classification, localization, and generative tasks; attribute and sentiment extraction for controlled blending; and end-to-end video–language models for dynamic metaphor captioning [2406.04886, 2310.10543].

## 3. Empirical Systems and Application Domains

HAIVMet encompasses a spectrum of systems attuned to distinct tasks:

- **Visual Metaphor Transfer and Blending**: Schema-driven frameworks (VMT) perform cross-domain logic transfer, preserving abstract relationships and style from a reference image onto novel targets, surpassing surface-level pixel remapping [2602.01335]. Creative Blends combines commonsense knowledge, semantic attribute matching, and generative diffusion to enable user-driven, semantically controlled visual blends [2502.16062].
- **Dialogue and Interaction Visualization**: Conversational DNA maps linguistic, affective, and topical features into interpretable visual metaphors (e.g., double helix), surfacing temporal interaction structure in both human–human and human–AI conversations [2508.07520].
- **Dream Narration and Affective Storytelling**: Metamorpheus scaffolds the co-creation of metaphorical visual scenes and text by guiding users through metaphor suggestion, prompt editing, and emotional arc construction; this supports reflection and personalized meaning-making [2403.00632].
- **Animation and Data Visualization**: DataSway empowers designers to produce coordinated metaphoric animations over SVG-based data visualizations via natural language interaction, VLM-assisted keyframing, group-wise motion coordination, and built-in data fidelity validation [2507.22051].
- **Visual and Multimodal Metaphor Understanding**: Benchmarks such as IRFL and MetaCLUE formalize multimodal figurative language recognition (detection, retrieval) and metaphor triplet extraction (concept, relation, concept), revealing the limitations of contemporaneous VL-PTMs and motivating modular, explainable pipelines [2303.15445, 2212.09898].
- **Video Metaphor Captioning**: Datasets and architectures (e.g., VMCD, GIT-LLaVA) support the captioning of metaphorical content in video, with creativity measured by metrics such as Average Concept Distance (ACD) [2406.04886].

## 4. Evaluation Protocols and Empirical Results

HAIVMet evaluation integrates automated, human, and hybrid metrics:

- **Automated Scores**: Metaphor Consistency (schema alignment), Analogy Appropriateness (attribute similarity), Conceptual Integration (blended harmony and novelty), CLIPScore, decomposition scores (\(r_{\mathrm{decomp}}\)), and meaning alignment (\(r_{\mathrm{MA}}\)) [2602.01335, 2508.18569].
- **Human Studies**: Protocols typically use Likert-scale ratings of recognizability, ingenuity, visual integration, and overall quality, with high inter-annotator agreement. Empirical results show HAIVMet-style pipelines outperform baselines (e.g., simple prompt transfer or single-agent generation) in metaphor fidelity, attribute fusion, and creativity, with improvements of +1 to +1.5 points on evaluative scales and 30–50% gains in automated metrics [2602.01335, 2305.14724].
- **Ablation and Error Analysis**: Studies dissect contributions of schema clarity, prompt design, and iterative refinement. Error modes in metaphor understanding include over-reliance on object co-occurrence, failure in abstract relation recognition, and hallucinations in explanatory tasks [2303.15445, 2405.01474].
- **Multimodal Generalization**: Video captioning models achieve highest ACDs when integrating pretraining on metaphoric image–text data (e.g., HAIVMet corpus) and employing template-driven captioning. Current approaches still display limited depth in semantic mapping [2406.04886].

## 5. Interpretability, Human Agency, and Theoretical Implications

HAIVMet systems foreground interpretability through explicit mapping and visualization of metaphor structure. Visualizations such as double helices, Sankey diagrams for attribute overlap, and concept–relation graphs expose otherwise latent narrative, emotional, or analogical architecture [2508.07520, 2502.16062]. User agency is promoted via editable forms, interactive storytelling, and prompt engineering interfaces, allowing co-creation rather than one-sided system output [2403.00632, 2507.22051].

The deployment of biological, perceptual, and literary metaphors as both technical and interactional scaffolds (e.g., the Library of Babel, The Aleph, Book of Sand) provokes reflection on the epistemic affordances and risks of infinite generativity, search, and agency in model–human collaboration [2402.07104, 2504.10101].

## 6. Limitations, Open Challenges, and Future Research

Several critical challenges and future directions are acknowledged:

- **Generalization to Cultural and Linguistic Diversity**: Most systems and benchmarks are English-only, introducing cultural bias and limiting abstraction to culturally embedded metaphors [2303.15445, 2212.09898].
- **Automated Reasoning over Abstract Relations**: Current VL models struggle with deep, non-literal mapping and often default to literal cue exploitation; progress requires fusion of reasoning modules, external knowledge bases, and explicit metaphoricity gating [2502.16062, 2303.15445].
- **Human-in-the-Loop Efficiency and Scalability**: Bottlenecks arise from expert curation and prompt engineering, suggesting a need for more efficient interactive interfaces and self-improving pipelines [2305.14724, 2508.18569].
- **Reward Design and Aesthetic Judgement**: Automated reward schemas integrating CLIPScore and BERTScore capture semantic alignment but not human aesthetic or emotional preference. Hybrid reward shaping with direct human feedback is an open area [2508.18569].
- **Compositional and Multimodal Extension**: Ongoing research targets multi-object blending, context-aware video and audio fusion, and joint conceptual–perceptual saccade visualization [2406.04886, 2504.10101].
- **Ethics and Safety**: As generative models become more powerful, considerations include fair representation, bias in analogy, inappropriate metaphors, and agency in infinite generation contexts [2402.07104, 2212.09898].

## 7. Synthesis and Cross-Domain Perspectives

HAIVMet represents an integrative evolution in human–AI creative collaboration, unifying structured analogy, deep learning, knowledge-backed inference, and interactive tooling. Advancements are tightly coupled to the emergence of modular, schema-aware agent frameworks, explainable evaluation protocols, and interactive systems that grant both computational rigor and human insight. Systems that can align arbitrary human intent with semantic, stylistic, and affective visual metaphor realization pave the way for new scientific, therapeutic, and artistic paradigms—enabling a future in which the architecture of ideas and affect can be visually composed, inspected, and evolved in partnership with intelligent machines [2508.07520, 2602.01335, 2305.14724, 2403.00632, 2212.09898].

---

**Key References:**  
- "Conversational DNA: A New Visual Language for Understanding Dialogue Structure in Human and AI" [2508.07520]  
- "Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning" [2602.01335]  
- "I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors" [2305.14724]  
- "The Human Visual System Can Inspire New Interaction Paradigms for LLMs" [2504.10101]  
- "The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation" [2508.18569]  
- "Metamorpheus: Interactive, Affective, and Creative Dream Narration Through Metaphorical Visual Storytelling" [2403.00632]  
- "DataSway: Vivifying Metaphoric Visualization with Animation Clip Generation and Coordination" [2507.22051]  
- "Creative Blends of Visual Concepts" [2502.16062]  
- "MetaCLUE: Towards Comprehensive Visual Metaphors Research" [2212.09898]

Source: https://www.emergentmind.com/topics/haivmet-human-ai-visual-metaphors