---
title: 'BioMetaphor: Biology-Inspired AI Models'
url: https://www.emergentmind.com/topics/biometaphor
type: topic
---

# BioMetaphor: Biology-Inspired AI Models

BioMetaphor denotes a family of research programs in which biological, bodily, or organismic structure is used to organize metaphor analysis, computational modeling, interface design, and system construction. In current usage, the term names both a specific human-centered GenAI framework for expressing biodata in virtual or hybrid co-present events and a broader line of inquiry into how embodiment-coded semantics, biomedical metaphor, Image Schemas, organismic analogies, and biologically inspired architectures shape artificial systems and human communication [2509.11600]. Taken together, these works suggest that BioMetaphor is less a single standardized doctrine than a heterogeneous research space spanning weak grounding through language, domain-specific metaphor resources, multimodal generation, and the transfer of biological concepts into engineering and theory building [2305.03445][2206.04603].

## 1. Conceptual scope and terminological boundaries

BioMetaphor research is organized by a recurring distinction between metaphor as heuristic language and metaphor as a basis for formal or operational design. Lester Ingber’s “Biological Impact on Military Intelligence: Application or Metaphor?” explicitly frames this distinction as a question of whether biological intelligence provides a bona fide application to military intelligence or only a pedagogically useful metaphor. In that paper, the biological source model is the neocortex as formalized by Statistical Mechanics of Neocortical Interactions, and the transfer target is Ideas by Statistical Mechanics for the propagation and evolution of ideas in populations. The argument is not that populations literally are brains, but that a validated multiscale stochastic formalism may be exported from neocortical modeling to idea propagation [1404.1360].

A closely related boundary appears in work on organismic analogies outside language technology. In urban science, the long-standing metaphor of “cities as organisms” is treated not as a literal ontological claim but as a heuristic and modeling tradition linking cities to evolution, morphogenesis, metabolism, autopoiesis, innovation, and collective intelligence. Raimbault’s citation-network study of around 225,000 papers argues that the most productive transfer occurs when such biological concepts are formalized through Artificial Life methods such as cellular automata, agent-based models, and morphogenetic approaches, rather than left as rhetoric [2002.12926]. In soft-matter engineering, “Meta-creatures” adopts an even more explicit organismal vocabulary—“omnipotent hydrogel cell,” “meta-nerve fibres,” “meta-skin,” “meta-cardiovascular system,” and “meta-creature”—but the paper itself is best read as a metaphor-heavy materials-and-systems study rather than a report of literal synthetic life [2408.08573].

The broadest philosophical version of BioMetaphor appears in the search for alternatives to the computer metaphor of mind and brain. “In search for an alternative to the computer metaphor of the mind and brain” does not propose one replacement image; it assembles a plurality of biologically grounded alternatives, including cascade dynamics, control systems, dissipative structures, complexity, radical embodied computation, resonance, and the brain as a fractal antenna. A central claim across those contributions is that the computer metaphor has become too narrow to capture embodied, context-sensitive, adaptive, goal-directed behavior in changing environments [2206.04603]. A plausible implication is that BioMetaphor often functions as an attempt to restore body, environment, development, and multi-scale organization to accounts of meaning and intelligence.

## 2. Embodiment, grounding, and metaphor processing in language models

One of the narrowest and most testable forms of BioMetaphor asks whether text-only language models are sensitive to bodily structure encoded in language. “LMs stand their Ground” defines embodiment cognitively and simulation-wise as the extent to which “the action can be simulated by a brain in a body,” then operationalizes it lexically using Sidhu et al.’s 1–7 verb ratings. On a derived Fig-QA subset, the embodied corpus \(C_{Emb}\) contains 1,438 phrases with at least one verb matched to the embodiment norms, while the matched-size control \(C_{NoE}\) also contains 1,438 items; \(C_{Liu}\) refers to the 1,146-item Fig-QA test set. The task is zero-shot forced-choice interpretation using the suffix “that is to say,” with candidate scoring by token log probabilities rather than free generation [2305.03445].

The main empirical result is modest but consistent. On \(C_{Emb}\), all larger model variants outperform their smaller paired counterparts in raw accuracy, and all larger versions above roughly 1B parameters show significant positive embodiment correlations: GPT-3 large accuracy \(0.667\), \(r=0.056\), \(p=0.034\); OPT 13B accuracy \(0.627\), \(r=0.056\), \(p=0.034\); GPT-NeoX 20B accuracy \(0.648\), \(r=0.073\), \(p=0.005\); GPT-2 XL 1.5B accuracy \(0.606\), \(r=0.069\), \(p=0.009\). GPT-3 small is an exception at reported ~350M size, with accuracy \(0.594\), \(r=0.062\), \(p=0.018\). The effect is small: all correlations are positive but below \(0.1\), and the paper explicitly states that the result is not that embodiment dominates metaphor understanding. Concreteness in context shows no significant correlation with model performance for any LM, and multicollinearity is limited, with VIF values AoA \(1.610\), frequency \(1.345\), embodiment \(1.326\), and word length \(1.017\). The authors therefore interpret the result as evidence for indirect grounding through linguistic distributions rather than genuine embodiment. Their distinction between “second-order grounding” and “first-order grounding” is central: human bodily experience has shaped metaphorical language, and the model learns from that language, but there is no body, no perception-action loop, and no independent experiential basis for source-domain mappings [2305.03445].

A later mechanistic study reframes the question from performance to internal organization. “Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy” introduces conditional scale entropy, defined over the wavelet-scaled residual-stream trajectory by
\[
H(\mathrm{scale} \mid b_k) = -\sum_{j=1}^{S} p(a_j \mid b_k)\log p(a_j \mid b_k),
\]
and proves that the quantity is invariant to overall update magnitude. Across GPT-2 Small, GPT-2 Medium, GPT-2 Large, LLaMA-2 7B, and GPT-oss 20B, metaphorical tokens produce significantly higher spectral breadth than literal tokens at contiguous positions, with the active zone recurring in an early-to-mid relative-depth range. The effect also converges with an independent analysis of 200 naturalistic VUA pairs and survives controls for semantic complexity and matched propositional content [2605.21391]. This suggests that BioMetaphor research on language models now spans both psycholinguistic proxy variables and post-hoc interpretability of depth-structured computation.

## 3. BioMetaphor as biodata representation for co-present events

In its most explicit naming, BioMetaphor is a human-centered generative-AI framework for turning biodata—especially heart rate (HR), respiration effort (RE), and galvanic skin response (GSR)—into context-appropriate visual social cues for virtual or hybrid co-present events such as VR galleries, sports matches, and concerts. The framework is positioned as an alternative or complement to avatar-centered social interaction, on the premise that large-scale collective settings often care less about exact face or gesture than about shared inner state, atmosphere, arousal, and synchrony. The paper’s central design claim is that biodata is not self-explanatory and therefore requires “technology-mediated biodata representation” rather than raw graphs or direct linear mappings [2509.11600].

The framework is grounded in a user-elicitation workshop with 30 HCI experts or postgraduate students in HCI-related fields. Ages ranged from 21 to 50, with mean 26.6 and SD 7.7; 18 participants were male and 12 female. Participants used a PICO 4 Pro headset and three text-to-image tools—Midjourney, DALL-E, and Stable Diffusion. After technical sensitization through a 360-degree DJ performance video with plain biodata visualizations, they experienced three 50-second VR scenarios: an art gallery, a table tennis match, and a virtual concert. The study retained 71 usable responses: 24 for gallery, 24 for sports, and 23 for concert. Visual modality dominated across all scenes: 88% visual in gallery, 79% in sports, and 87% in concert. Mean preferred “level of biodata interpretation” was 54.75 (SD 27.92), and mean preferred “level of biodata understanding” was 33.82 (SD 25.03). Using Lakoff and Johnson’s taxonomy, ontological metaphors were most common at 52.1% of responses, structural metaphors accounted for 22.5%, orientational metaphors 11.3%, and plain data visualization, sonification, or haptification 28.2%, often alongside metaphor [2509.11600].

These findings are operationalized in four modules. The Biodata Modelling Module converts biosensory input into a normalized valence-arousal pair using Russell’s circumplex model; the paper gives \((0.14, 0.85)\) as an example. The LLM-driven Metaphor Building Module then performs four chain-of-thought steps: Inner State Inference, Metaphor Construction, Event Adaptation, and Prompt Generation. The Representation Generation Module uses Stable Diffusion, specifically SDXL plus LoRA, to create a panoramic VR-ready image. The Virtual Integration Module renders the result in Unity as a skybox texture and streams it to an HMD such as the PICO 4 Pro [2509.11600].

The demonstration evaluates eight prototypical valence-arousal pairs evenly spaced around the circumplex at 45-degree intervals across three scenarios, using GPT-4o, DeepSeek-Chat, and one image generator, for 48 outcomes total. The authors report that both LLMs generally infer emotional ranges aligned with the circumplex model and can combine metaphor types in event-situated prompts, but structural metaphors are often oversimplified and SDXL sometimes fails to realize metaphorical details. The paper therefore presents BioMetaphor as a feasibility test rather than a controlled end-user evaluation [2509.11600]. A common misconception addressed by the framework is that biodata representation is mainly a sensing problem; the paper explicitly reframes it as a cognitive-design problem.

## 4. Biomedical metaphor corpora, repositories, and extraction workflows

A major strand of BioMetaphor research concerns the construction of biomedical metaphor resources. The Medical Metaphors Corpus (MCC) is introduced as “the first openly released annotated resource dedicated to metaphorical language across medical and biological discourse” and contains 792 annotated scientific conceptual metaphors from nine sources spanning literature, news, social media, interviews, and crowdsourcing. The corpus includes 82 distinct metaphor types, 24 unique target domains, and 38 unique source domains. Each instance stores a binary metaphoricity label, a graded metaphoricity score on a 0–7 scale, source–target conceptual mappings, and provenance metadata. The annotator pool consists of 27 advanced students in Informatica Umanistica and 15 online linguists, for 42 annotators total; each sentence received at least two annotations. Inter-annotator agreement is intentionally moderate: Fleiss’ kappa \(= 0.23\), average percent agreement \(= 60\%\), and average Pearson \(r = 0.4\) with Spearman \(\rho = 0.4\) for graded ratings. Across all annotations, 353 items were labeled yes, 305 no, and 134 ties [2508.07993].

MCC also supplies a biomedical LLM benchmark. The evaluated models are GPT-4, o1-preview, o3-mini, DeepSeek, and Claude Opus 4, all in zero-shot settings. Performance is characterized as modest, with o1-preview best overall at accuracy \(0.716\) and weighted accuracy \(0.758\). The paper reports that all models improve under consensus weighting by roughly 3–4.6%, indicating better behavior on clearer cases than on ambiguous ones, and identifies a conservative precision–recall profile in which models under-detect metaphor unless cues are overt [2508.07993]. The benchmark’s importance for BioMetaphor lies in its focus on highly conventionalized scientific metaphor, where terms such as “invasion” or “transport” may function simultaneously as technical vocabulary and as metaphorical inheritances.

A complementary workflow is reported for Dutch oncology discourse. “Dutch Metaphor Extraction from Cancer Patients’ Interviews and Forum Data using LLMs and Human in the Loop” compiles the HealthQuote.NL corpus from 13 interview transcripts and a pilot sample of the first 100 cancer forum blog posts. After human verification, the paper reports 130 total metaphors: 65 from interview data and 65 from forum data. The workflow uses chunking with maximum token length 4000, overlap 40 tokens, context window 32768, and prompt strategies including chain of thought, few-shot learning, and self-prompting; human filtering removes hallucinations, paraphrases, over-interpretations, and generic idioms. The paper does not report precision, recall, F1, or inter-annotator agreement, so its contribution is better understood as exploratory methodology and corpus building than as a benchmark [2511.06427].

Infrastructure for sharing such resources is provided by MetaphorShare, described as “a dynamic collaborative repository of open metaphor datasets.” The platform supports upload, download, search, and label functionalities, with a CSV-based upload format requiring a `tagged_sentence` column marked by inline XML-like tags `<m: metaphoric, l: literal, t: target, u: free tag >`. Search is backed by Elasticsearch, metadata are stored in PostgreSQL, and the site currently integrates 25 datasets, though most are in English. For BioMetaphor projects, MetaphorShare is relevant not because it already centers biomedical datasets, but because it offers a repository model for normalization, search, and redistribution of open metaphor corpora [2411.18260]. This suggests an infrastructural turn in which BioMetaphor becomes not only a theoretical issue but also a data-engineering and curation problem.

## 5. Multimodal enactment: gestures and metaphorical communication spaces

BioMetaphor also appears in work where metaphor is enacted bodily or interactively rather than only detected in text. META4 addresses metaphoric gesture generation by inserting Image Schemas into a speech-driven motion pipeline. The system first predicts one of 14 Image Schema classes—CENTER-PERIPHERY, CONTACT, CONTAINMENT, COVERING, FORCE, LINK, OBJECT, PART-WHOLE, SCALE, SOURCE_PATH_GOAL, SPLITTING, SUBSTANCE, SUPPORT, and VERTICALITY—from text using a fine-tuned BERT Base Cased classifier called BERTIS, trained on 1,994 English utterance samples. It then combines the schema representation with an Audio Spectrogram Transformer embedding and conditions a Transformer decoder to generate 2D upper-body poses over 64-frame segments. The core equations are
\[
h^i_{IS} = E_{BERTIS}(X_{text}), \qquad
h_{audio} = E_{audio}(X_{audio}),
\]
\[
h_{Audio\_IS} = [E_{audio}(X_{audio}), E_{BERTIS}(X_{text})], \qquad
\widehat{Y}_{Pose} = G_{pose}(h_{Audio\_IS}).
\]
BERTIS reaches overall accuracy \(0.93\). In gesture generation, the full model outperforms the Image-Schema ablation in both seen and unseen speaker settings: for seen speakers, RMSE improves from \(0.02004\) to \(0.01627\); for unseen speakers, from \(0.02520\) to \(0.02100\) [2311.05481]. The paper’s importance for BioMetaphor is that it operationalizes embodied semantic structure as a control variable rather than relying on speech timing alone.

A different multimodal extension appears in MetaphorChat, a “metaphorical chatting space” for expressing and understanding inner feelings. Although the system is not explicitly framed as biological metaphor, it demonstrates a cognate design principle: metaphor is implemented across graphics, sound, music, movement, interaction rules, and synchronous chat, rather than remaining a textual figure. The two scenes, “On a Boat” and “On a Train,” were developed through autobiographical design to support communication about uncertainty, insecurity, being lost, hopelessness, and the transience of interpersonal relationships. The design process is summarized in three stages—Metaphor Mapping, Integrating metaphors into the chat system, and Tweaking the “vibe”—and includes interactions such as only the sharer being able to move the boat and the train moving only after saying goodbye. Qualitative self-use suggests that these shared metaphorical environments can lower the threshold for discussing difficult feelings and make abstract states more understandable to others [2502.07125]. A plausible implication is that BioMetaphor in HCI often depends on enactment: the metaphor must be inhabited, not merely named.

## 6. Biological metaphor as architectural and synthetic design principle

In AI architecture, BioMetaphor can function as a meta-architectural proposal. Alicea and Parent’s “Meta-brain Models: biologically-inspired cognitive agents” argues that neither neural networks alone nor purely symbolic systems are sufficient for richer cognition. Their answer is a layered embodied hybrid model composed of representation-free or representation-poor lower layers, sparse intermediate layers, and representation-rich higher layers, with feedforward and feedback connectivity inspired by the neocortical-thalamic relationship. The framework also includes an innate basement layer with mutational and transcription sublayers, likened to genome and epigenetic expression, and ties cognition to embodiment, morphogenesis, and development. The paper is largely conceptual and does not provide formal mathematical notation or update laws, but it explicitly proposes layered heterogeneity, anatomical explicitness, and developmental organization as principles for neuro-symbolic AI [2109.11938].

In soft robotic and materials terms, “Meta-creatures” translates developmental and organismal language into hydrogel engineering. The smallest unit is the “omnipotent hydrogel cell,” arranged as \(\mathrm{H | (rGO)C | L | (rGO)A | H}\). For a 25 mm OHC with cross-sectional area \(0.785\ \mathrm{cm}^2\), the paper reports currents from \(-31.11\) to \(+15.37\ \mu\mathrm{A}\) and voltages from \(-53.55\) to \(+15.39\ \mathrm{mV}\), with a relatively stable voltage of \(14.68\ \mathrm{mV}\) for 5.46 h and stable current of \(13.58\ \mu\mathrm{A}\) for 2.51 h. These modules are then transformed into meta-nerve fibres, a meta-mouth, meta-skin, a meta-cardiovascular system, and a bird-like “meta-creature.” The paper explicitly uses terms such as proliferation, regeneration, metabolism, and reconfiguration, but it also makes clear that these are engineered hydrogel analogues with partial functional similarities to living systems rather than literal organisms [2408.08573].

The same architectural use of BioMetaphor appears at urban scale. Raimbault’s “Cities as they could be” reconstructs the research landscape linking Artificial Life, AI, and urban systems through a citation network with \(|V| = 224{,}510\) papers and \(|E| = 315{,}829\) citation links, with a core network of \(|V| = 48{,}657\) and \(|E| = 139{,}931\) and directed modularity \(0.84\). The paper argues that biological imports such as morphogenesis, urban metabolism, autopoiesis, innovation, and open-ended evolution are most useful when they become generative theories of urban form and function rather than decorative analogy [2002.12926]. Across these examples, BioMetaphor operates as a design language that organizes modules, scales, developmental stages, and feedback relations.

## 7. Debates, limitations, and research directions

A recurring debate in BioMetaphor concerns the difference between indirect biological grounding and literal biological equivalence. In language-model work, the strongest positive claims are deliberately weak: larger text-only models appear to recover embodied regularities from linguistic distributions, but this remains “second-order grounding” rather than full biological embodiment, with no body, no sensorimotor contingencies, and no premotor or perceptual simulation analogous to human processing [2305.03445]. In synthetic-systems work, “meta-creature” language similarly exceeds literal biology: the hydrogel platform exhibits bio-inspired bioelectricity and partial organ-like functions, but not autonomous organismality in the biological sense [2408.08573].

Methodological limitations are equally prominent. The BioMetaphor biodata framework is based on 30 HCI experts and a prototype demonstration rather than a controlled user evaluation; it depends heavily on prompt engineering, and the authors explicitly note that structural metaphors are often oversimplified and image synthesis can fail to realize intended metaphorical details. The same paper foregrounds privacy and cultural sensitivity because biodata is intimate and AI-generated metaphors can carry socially inappropriate meanings [2509.11600]. MCC is English-only, has moderate inter-annotator agreement, no train/dev/test split, and evaluates only zero-shot frontier LLMs rather than fine-tuned biomedical systems [2508.07993]. HealthQuote.NL relies on human verification without a gold-standard benchmark or reported inter-annotator agreement, and privacy constraints require paraphrased or synthetic forum examples rather than original patient text [2511.06427]. MetaphorShare is designed for open datasets, which makes it useful for public biomedical or health-communication corpora but less suited to restricted clinical text unless it is fully de-identified and legally shareable [2411.18260]. META4 lacks subjective evaluation on embodied conversational agents, omits explicit loss functions, and uses 2D upper-body poses without fingers, even though hand shape is often crucial for metaphoric gesture semantics [2311.05481].

Future research directions are correspondingly diverse. The embodiment study proposes testing other figurative datasets such as FLUTE and IMPLI, examining embodiment effects on non-figurative BIG-bench tasks, and expanding beyond Fig-QA [2305.03445]. The biodata framework calls for comparing prompt strategies, moving beyond prompts to fine-tuning, and validating whether users actually experience stronger empathy, co-presence, or understanding [2509.11600]. MCC proposes fine-tuning or continued pre-training on biomedical metaphor, adding richer annotations such as emotional valence and explanatory clarity, integrating symbolic ontologies with LLMs, and extending to other scientific domains and languages [2508.07993]. MetaphorShare plans an online annotation tool and semi-automatic labeling workflows [2411.18260]. The metaphor-processing interpretability work proposes continuously rated metaphor novelty, causal intervention via scale-channel suppression, and extension from metaphor to irony, metonymy, polysemy, and multi-step reasoning [2605.21391]. Taken together, these directions suggest that BioMetaphor is moving from rhetorical borrowing toward empirically constrained, multimodal, and infrastructurally supported research programs.

Source: https://www.emergentmind.com/topics/biometaphor