Papers
Topics
Authors
Recent
Search
2000 character limit reached

HAIVMet: Human–AI Visual Metaphors

Updated 17 March 2026
  • HAIVMet is a computational paradigm that leverages human-AI collaboration to generate, analyze, and explain nonliteral meanings in visual metaphors.
  • The framework employs schema extraction, agent-based reasoning, and prompt decomposition to reframe abstract concepts through visual modalities.
  • Empirical results demonstrate improved metaphor fidelity and creative visual synthesis, bridging abstract ideas with concrete imagery.

Human–AI Visual Metaphors (HAIVMet) designate a computational paradigm and family of systems that leverage the complementary interpretive powers of humans and artificial intelligence to generate, analyze, transfer, or explain nonliteral meaning in visual modalities. Rooted in cognitive metaphor theory and extended through advances in generative modeling, schema reasoning, and collaborative design interfaces, HAIVMet reframes both the construction and interpretation of visual metaphors, enabling richer, more actionable connections between abstract concepts and concrete imagery. Contemporary HAIVMet systems integrate LLMs, vision–LLMs, and interactive agents with human-in-the-loop workflows across disparate domains, including dialogue visualization, creative image synthesis, video captioning, scientific visualization, and affective design (Lin, 11 Aug 2025, Xu et al., 1 Feb 2026, Chakrabarty et al., 2023, Robinson et al., 14 Apr 2025, Koushik et al., 26 Aug 2025, Wan et al., 2024, Xie et al., 29 Jul 2025).

1. Theoretical Foundations and Metaphorical Mapping

HAIVMet instantiates foundational principles from cognitive linguistics and visual rhetoric by operationalizing core constructs such as cross-domain mapping, conceptual blending, and schema invariance. Systems in this class encode the abstract structure of a metaphor—typically as a mapping between source and target domains or as a triplet (primary concept, relation, secondary concept)—and instantiate this mapping in visual modalities via algorithmic or agentic means (Akula et al., 2022, Sun et al., 22 Feb 2025, Yosef et al., 2023).

For example, in the schema-driven VMT (Visual Metaphor Transfer), a reference image is abstracted into a universal schema grammar encoding entities, generic-space relations (e.g., ontological, functional), aesthetic attributes, and violation points. The system then recomposes the abstract logic onto a distinct subject, ensuring the invariant relational mapping persists (Xu et al., 1 Feb 2026). Analogously, in conversational visualization (Conversational DNA), linguistic features such as contribution volume, emotional valence, and semantic alignment are recast as intertwined helical strands, making the architecture of dialogue “seeable” as a living biological metaphor (Lin, 11 Aug 2025).

2. Computational and Algorithmic Frameworks

HAIVMet workflows blend symbolic reasoning, neural generation, and joint human–AI curation. Central methodologies include:

  • Schema Extraction and Grammar Encoding: Systems abstract high-level logic from visual or textual input using parsers, LLMs, and CLIP-embedding pipelines to identify component roles, attributes, and analogical invariants (Xu et al., 1 Feb 2026, Sun et al., 22 Feb 2025, Akula et al., 2022).
  • Agent-based Reasoning: The VMT framework allocates specialized agents for perception (schema extraction), transfer (invariant preservation and adaptation), generation (prompt synthesis), and hierarchical diagnostics. Interaction between these agents enables closed-loop correction and robust metaphor reification (Xu et al., 1 Feb 2026).
  • Prompt Decomposition: Structured prompting and S–T–M (Source–Target–Meaning) frameworks force explicit delineation of metaphorical structure, which is then injected into prompt engineering for text-to-image models (Koushik et al., 26 Aug 2025, Chakrabarty et al., 2023). LaTeX-formatted reward functions (R(I,P,D)\mathcal{R}(I, P, D)) aggregate decomposition quality, CLIPScore, and meaning-alignment metrics to guide prompt refinement.
  • Human–AI Collaboration: HAIVMet systems typically embed human review/validation steps. These include chain-of-thought (CoT) prompting for visual elaboration followed by expert or crowd-based filtering, as in the HAIVMet dataset (Chakrabarty et al., 2023), collaborative metaphor suggestion and editing in affective interfaces (Wan et al., 2024), or interactive annotation and feedback in understanding pipelines (Yosef et al., 2023, Akula et al., 2022).
  • Joint Multimodal Reasoning: Architectures span vision–language transformers for classification, localization, and generative tasks; attribute and sentiment extraction for controlled blending; and end-to-end video–LLMs for dynamic metaphor captioning (Kalarani et al., 2024, Shahmohammadi et al., 2023).

3. Empirical Systems and Application Domains

HAIVMet encompasses a spectrum of systems attuned to distinct tasks:

  • Visual Metaphor Transfer and Blending: Schema-driven frameworks (VMT) perform cross-domain logic transfer, preserving abstract relationships and style from a reference image onto novel targets, surpassing surface-level pixel remapping (Xu et al., 1 Feb 2026). Creative Blends combines commonsense knowledge, semantic attribute matching, and generative diffusion to enable user-driven, semantically controlled visual blends (Sun et al., 22 Feb 2025).
  • Dialogue and Interaction Visualization: Conversational DNA maps linguistic, affective, and topical features into interpretable visual metaphors (e.g., double helix), surfacing temporal interaction structure in both human–human and human–AI conversations (Lin, 11 Aug 2025).
  • Dream Narration and Affective Storytelling: Metamorpheus scaffolds the co-creation of metaphorical visual scenes and text by guiding users through metaphor suggestion, prompt editing, and emotional arc construction; this supports reflection and personalized meaning-making (Wan et al., 2024).
  • Animation and Data Visualization: DataSway empowers designers to produce coordinated metaphoric animations over SVG-based data visualizations via natural language interaction, VLM-assisted keyframing, group-wise motion coordination, and built-in data fidelity validation (Xie et al., 29 Jul 2025).
  • Visual and Multimodal Metaphor Understanding: Benchmarks such as IRFL and MetaCLUE formalize multimodal figurative language recognition (detection, retrieval) and metaphor triplet extraction (concept, relation, concept), revealing the limitations of contemporaneous VL-PTMs and motivating modular, explainable pipelines (Yosef et al., 2023, Akula et al., 2022).
  • Video Metaphor Captioning: Datasets and architectures (e.g., VMCD, GIT-LLaVA) support the captioning of metaphorical content in video, with creativity measured by metrics such as Average Concept Distance (ACD) (Kalarani et al., 2024).

4. Evaluation Protocols and Empirical Results

HAIVMet evaluation integrates automated, human, and hybrid metrics:

  • Automated Scores: Metaphor Consistency (schema alignment), Analogy Appropriateness (attribute similarity), Conceptual Integration (blended harmony and novelty), CLIPScore, decomposition scores (rdecompr_{\mathrm{decomp}}), and meaning alignment (rMAr_{\mathrm{MA}}) (Xu et al., 1 Feb 2026, Koushik et al., 26 Aug 2025).
  • Human Studies: Protocols typically use Likert-scale ratings of recognizability, ingenuity, visual integration, and overall quality, with high inter-annotator agreement. Empirical results show HAIVMet-style pipelines outperform baselines (e.g., simple prompt transfer or single-agent generation) in metaphor fidelity, attribute fusion, and creativity, with improvements of +1 to +1.5 points on evaluative scales and 30–50% gains in automated metrics (Xu et al., 1 Feb 2026, Chakrabarty et al., 2023).
  • Ablation and Error Analysis: Studies dissect contributions of schema clarity, prompt design, and iterative refinement. Error modes in metaphor understanding include over-reliance on object co-occurrence, failure in abstract relation recognition, and hallucinations in explanatory tasks (Yosef et al., 2023, Saakyan et al., 2024).
  • Multimodal Generalization: Video captioning models achieve highest ACDs when integrating pretraining on metaphoric image–text data (e.g., HAIVMet corpus) and employing template-driven captioning. Current approaches still display limited depth in semantic mapping (Kalarani et al., 2024).

5. Interpretability, Human Agency, and Theoretical Implications

HAIVMet systems foreground interpretability through explicit mapping and visualization of metaphor structure. Visualizations such as double helices, Sankey diagrams for attribute overlap, and concept–relation graphs expose otherwise latent narrative, emotional, or analogical architecture (Lin, 11 Aug 2025, Sun et al., 22 Feb 2025). User agency is promoted via editable forms, interactive storytelling, and prompt engineering interfaces, allowing co-creation rather than one-sided system output (Wan et al., 2024, Xie et al., 29 Jul 2025).

The deployment of biological, perceptual, and literary metaphors as both technical and interactional scaffolds (e.g., the Library of Babel, The Aleph, Book of Sand) provokes reflection on the epistemic affordances and risks of infinite generativity, search, and agency in model–human collaboration (Ramos et al., 2024, Robinson et al., 14 Apr 2025).

6. Limitations, Open Challenges, and Future Research

Several critical challenges and future directions are acknowledged:

  • Generalization to Cultural and Linguistic Diversity: Most systems and benchmarks are English-only, introducing cultural bias and limiting abstraction to culturally embedded metaphors (Yosef et al., 2023, Akula et al., 2022).
  • Automated Reasoning over Abstract Relations: Current VL models struggle with deep, non-literal mapping and often default to literal cue exploitation; progress requires fusion of reasoning modules, external knowledge bases, and explicit metaphoricity gating (Sun et al., 22 Feb 2025, Yosef et al., 2023).
  • Human-in-the-Loop Efficiency and Scalability: Bottlenecks arise from expert curation and prompt engineering, suggesting a need for more efficient interactive interfaces and self-improving pipelines (Chakrabarty et al., 2023, Koushik et al., 26 Aug 2025).
  • Reward Design and Aesthetic Judgement: Automated reward schemas integrating CLIPScore and BERTScore capture semantic alignment but not human aesthetic or emotional preference. Hybrid reward shaping with direct human feedback is an open area (Koushik et al., 26 Aug 2025).
  • Compositional and Multimodal Extension: Ongoing research targets multi-object blending, context-aware video and audio fusion, and joint conceptual–perceptual saccade visualization (Kalarani et al., 2024, Robinson et al., 14 Apr 2025).
  • Ethics and Safety: As generative models become more powerful, considerations include fair representation, bias in analogy, inappropriate metaphors, and agency in infinite generation contexts (Ramos et al., 2024, Akula et al., 2022).

7. Synthesis and Cross-Domain Perspectives

HAIVMet represents an integrative evolution in human–AI creative collaboration, unifying structured analogy, deep learning, knowledge-backed inference, and interactive tooling. Advancements are tightly coupled to the emergence of modular, schema-aware agent frameworks, explainable evaluation protocols, and interactive systems that grant both computational rigor and human insight. Systems that can align arbitrary human intent with semantic, stylistic, and affective visual metaphor realization pave the way for new scientific, therapeutic, and artistic paradigms—enabling a future in which the architecture of ideas and affect can be visually composed, inspected, and evolved in partnership with intelligent machines (Lin, 11 Aug 2025, Xu et al., 1 Feb 2026, Chakrabarty et al., 2023, Wan et al., 2024, Akula et al., 2022).


Key References:

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HAIVMet (Human-AI Visual Metaphors).