---
title: Imaginative Perception
url: https://www.emergentmind.com/topics/imaginative-perception
type: topic
---

# Imaginative Perception

Imaginative perception denotes forms of perception in which what is apprehended is not exhausted by the immediately given stimulus, but is organized through metaphor, associative completion, internal simulation, or generative modeling into scenes, relations, and possibilities that are absent, hidden, future, or purely mental. In contemporary research, the term spans several technically distinct phenomena: making inner life visually manipulable through light-based metaphors, representing “what it is like to see” as directions in latent space, decoding percept-like imagery from brain activity, synthesizing unseen viewpoints for robots, and using scientific diagrams as props for visual imagination [1501.00029] [2012.14283] [2404.05468] [2505.08340].

## 1. Philosophical scope and conceptual lineage

A recurring claim across the literature is that perception is not merely registration of sensory data but a constructive act that recruits prior structure. In “Latent Compass,” perception is described as “the vehicle for generalization,” with “prickliness” functioning as a “concrete abstraction” that can move across modalities, from lemons to corners to sight [2012.14283]. In philosophy of scientific imagery, Meynell’s view of images as aids to visual imagination is used to explain how readers “see through” a diagram to a spatial–causal astrophysical scenario, rather than merely reading marks on a page [2505.08340]. In “From Virtual Reality to the Emerging Discipline of Perception Engineering,” this constructive dimension is generalized into a producer–receiver framework in which one agent alters the environment with the intent to alter another agent’s perceptual experience [2403.18588].

A more radical formulation appears in “Illusions - a model of mind,” where perception, imagination, and thought are treated as unfree and passive outcomes of associative neural processes rather than acts of a central inner observer [1707.09379]. On that view, the experienced world, the sense of an “I,” and even the intuition of free choice are themselves products of constructive processing. This suggests a broad encyclopedia-level definition: imaginative perception is perception insofar as it depends on internal structure that exceeds literal stimulus content, whether that structure is phenomenological, symbolic, neural, or computational.

## 2. Representational forms and mediating structures

One major research line externalizes imaginative perception through explicit representational media. LIVEIA turns psychological states and relationships into an immersive optical language: ambient light represents creative life force; psyches are spheres that amplify and transform light; shell thickness and opacity represent guardedness or openness; and beams of light represent thoughts and feelings with intensity, diffuseness, direction, and waveform-like qualitative character [1501.00029]. The paper explicitly proposes a conceptual mapping $L : P \rightarrow V$ from psychological states to visual parameters, making inner life perceptible without treating it as a measurement task.

Another line operationalizes imagination as structured expansion in conceptual space. “From Knowledge Map to Mind Map: Artificial Imagination” defines artificial imagination as more than semantic expansion: it includes lexical similarity via longest common substring, phonological similarity via homophony, an artist-specific knowledge network with 12,000 entities and 1,754 relations, and Dadaist cross-domain or random associations [1903.01080]. The resulting “Mind Map” is neither purely taxonomic nor purely embedding-based; it is a visual-conceptual network in which a seed word is transformed into a stylized Shan Shui scene through domain classification into Architecture, Mountain, River, Grassland, Road, and Lake [1903.01080].

A third representational form appears in “Imaginative Perception Tokens,” where imaginative perception is encoded as intermediate perceptual representations corresponding to what a multimodal model would perceive under alternative spatial configurations [2606.03988]. Here the imagined state is neither free text nor a final output image alone; it is a supervised latent image representation aligned with unobserved viewpoints, path midpoints, or bird’s-eye spatial summaries. The shared pattern across these systems is that imaginative perception is not treated as ineffable: it is externalized as manipulable structure.

## 3. Generative mechanisms and uncertainty-guided imagination

Several computational systems treat imaginative perception as conditional generation over missing or future perceptual states. “Visual Attention in Imaginative Agents” gives a recurrent agent discrete fixations, a generative model over full scenes, and a policy that plans the next fixation from uncertainty in imagined scenes [2104.00177]. At time $t$, the fixation history is summarized by an RNN state $h_t$; samples from a conditional latent distribution yield multiple plausible full-scene hypotheses; and a variance map,
$$
V_t = \mathbb{E}_{z_k \sim p(z_k \mid h_t)} \left\| D(f_{h_t}(z_k)) - \mu_t \right\|^2,
$$
guides the next fixation toward regions where those hypotheses disagree most [2104.00177]. Imaginative perception here is explicitly probabilistic and active.

“Imagination-enabled Robot Perception” instantiates a related principle for manipulation: the robot maintains a scene graph in an Artificial World, renders expected sensor data $S_{AW}$ from hypothesized object models, and compares it with real sensor data $S_{RW}$ while rejecting physically impossible scenes through physics simulation [2011.11397]. The comparison uses color histograms and the Hellinger distance, and the whole loop behaves like approximate inference over scene hypotheses. A plausible implication is that imaginative perception in robotics becomes strongest when visual generation is coupled to explicit physical plausibility.

“Mind-to-Image” extends the generative idea to mental imagery itself. Using about 6 hours of fMRI scans and a modified MindEye pipeline, it maps beta-weight summaries from regions including visual cortex, fusiform gyrus, temporal associative areas, and prefrontal cortex into CLIP and VAE latent spaces, then reconstructs images from weak imagination and pure imagination [2404.05468]. The reconstructions are imperfect, but the system reaches 91% portrait-versus-landscape classification accuracy for weak imagination and 88% for strong imagination, showing that internally generated imagery can be projected into a generative visual manifold [2404.05468].

## 4. Neural, cognitive, and phenomenological grounding

Mental imagery is treated in cognitive neuroscience as “percept-like experiences in the absence of sensory input.” “Predicting the imagined contents using brain activation” tests this directly with imagined monetary rewards and scrambled pictures: a support vector machine trained only on visually presented reward trials predicts whether subjects imagined a monetary reward or a scrambled picture with 75% accuracy, and the ROC analysis gives an AUC of 0.78 [2106.07355]. The relevant signal lies in dopaminergic midbrain regions, showing that imagined reward and perceived reward recruit overlapping neural structure [2106.07355].

The imagination/perception relation becomes richer when imagery is not merely recall. In “Mind-to-Image,” weak imagination activates basal temporal associative areas and fusiform regions, whereas strong imagination recruits broader prefrontal and temporal networks, suggesting a shift from memory-supported reactivation to more controlled constructive generation [2404.05468]. The result does not collapse imagination into perception; rather, it supports a family resemblance in representational format with a difference in control structure and cortical distribution.

Phenomenological accounts align with this pattern without reducing to it. “Latent Compass” takes Cézanne and the post-impressionists as paradigmatic figures for rendering not what was seen but “what it is like to see,” then reifies such experiential qualities as latent directions $z' = z + \lambda d$ in a GAN space [2012.14283]. “Illusions - a model of mind,” by contrast, argues that the sense of an observer and the character of perceptual experience arise from passive associative processes and silent inner communication [1707.09379]. Taken together, these works suggest that imaginative perception is neither mere fantasy nor ordinary sensation, but a regime of percept-like organization with heavy top-down participation.

## 5. Language, art, and multimodal AI

In multimodal AI, imaginative perception often appears as a bridge from underdetermined input to richer internal state. “Imaginations of WALL-E” makes this explicit: a text input $x$ is first transformed into a generated image $I = g(x; \theta_{\text{T2I}})$ by Stable Diffusion v2, then passed with the text into InstructBLIP for downstream interpretation [2308.10354]. In zero-shot evaluation, the imagination-inspired system reaches Weighted F1 scores of 46.74% on MELD and 25.23% on IEMOCAP, and an Overall F1 of 17.0% on CoQA, compared to 22.89%, 12.28%, and 7.0% from the strongest unimodal baselines reported there [2308.10354]. The system’s “Interpretable Misunderstanding” names divergences between human expectation and model reconstruction that remain semantically anchored.

A related but more explicitly grounded approach appears in “Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games.” There, an imagination module based on Regularized Auto-Encoders learns latent object embeddings directly from visual crops and scene context, replacing gold category labels at inference [2011.02917]. It improves zero-shot gameplay accuracy in CompGuessWhat?! by 8.26%, and boosts Oracle and Guesser accuracy by 2.08% and 12.86% on GuessWhat?! when no gold categories are available at inference time [2011.02917]. Imaginative perception in this setting is the ability to construct a concept-like latent from perception alone and use it dialogically.

Art-oriented systems make the same move with different media. The “Mind Map” system raises Dada-style expansions from 4.40% to 22.86% and author-style expansions from 2.42% to 15.61%, while human experts rate it higher than a semantic-only baseline on Linguistic Connection, Dadaism, and Overall Impression [1903.01080]. “Latent Compass” instead keeps creators in the loop by asking them to sort generated images into two groups, then training a linear SVM whose normal vector becomes a direction in latent space [2012.14283]. Both systems treat imaginative perception as structured deviation from literal semantic relevance, but in one case the medium is lexical-artistic expansion and in the other latent navigation.

## 6. Embodied and spatial imagination

Embodied agents provide the clearest contemporary testbed for imaginative perception as unseen-view reasoning. “Imagination at Inference” synthesizes missing in-hand views from external agent views via LoRA-finetuned ZeroNVS, then feeds those views into visuomotor policies at deployment [2509.15717]. In real-world strawberry picking with a Unitree Z1 arm, the success rates are 3/10 for agent view only, 8/10 with a real in-hand camera, and 6/10 with synthesized in-hand views, showing substantial recovery of performance without installing the physical wrist camera [2509.15717].

Astra generalizes the same idea into agentic world-model use. Astra-WM improves simulator-augmented Gemini-3-Flash on MMSI-Bench from 45.1 to 49.5, while Astra-VL improves the Qwen3-VL backbone from 29.8 to 38.8 on MMSI-Bench and from 36.8 to 42.7 on MindCube [2606.06476]. The crucial claim is not simply that imagined observations help, but that effective use requires learning when, where, and how to imagine. This makes imaginative perception a control problem over internal evidence acquisition rather than a passive hallucination layer.

Navigation systems reach similar conclusions. SALI equips a vision-and-language navigation agent with a reality–imagination hybrid memory and recurrent imagination tree, and reports state-of-the-art SPL while imagining high-fidelity RGB future scenes [2412.01857]. DreamNav, in zero-shot VLN-CE, couples egocentric view correction, trajectory-level planning, and an Imagination Predictor, outperforming the strongest egocentric baseline with extra information by up to 7.49% in SR and 18.15% in SPL [2509.11197]. A parallel line replaces explicit image generation with supervised intermediate states: Imaginative Perception Tokens improve Multiview Counting accuracy by 3.4% and can outperform textual chain-of-thought training, which the paper reports can substantially degrade performance under modality mismatch [2606.03988]. Across these systems, imaginative perception functions as a spatially grounded estimate of what would be seen under an alternative trajectory or viewpoint.

## 7. Scientific imagery, evaluation, and open problems

Imaginative perception is not confined to AI. In astrophysics, the Stellar Graveyard plot is analyzed as an image that may be sub-optimal for statistical argument yet valuable as an aid to visual imagination: circles and arrows prompt readers to imagine compact objects, merger dynamics, and causal structure that are not literally shown on the page [2505.08340]. The same paper extends this analysis to black-hole merger diagrams, Van den Heuvel binary evolution schematics, and supernova sketches, arguing that such images help readers grasp spatial–causal relations and counterfactual variation [2505.08340]. Scientific understanding, on this account, often requires imaginative perception rather than proposition-only reasoning.

Evaluation of imaginative perception remains heterogeneous. “What to Ask Next?” introduces TurtleSoup-Bench, a bilingual interactive benchmark of 800 puzzles, together with Mosaic-Agent and a three-part evaluation protocol measuring logical consistency, detail completion, and conclusion alignment [2508.10358]. “Classifying informative and imaginative prose using complex networks” shows that local topological measurements over function-word networks can discriminate informative from imaginative prose with up to 95% accuracy, with symmetry and accessibility among the most relevant features [1507.07826]. “From Knowledge Map to Mind Map” uses Relevance, Linguistic Connection, Dadaism, and Overall Impression as human-centric scores [1903.01080]. These evaluation regimes are not commensurable, but they converge on a common difficulty: imaginative perception is useful partly because it is structured and partly because it is open-ended.

The main controversies follow directly from that duality. Metaphorical systems such as LIVEIA face tensions between systematic mappings and idiosyncratic interpretation, along with early-stage implementation and limited empirical validation [1501.00029]. Generative systems inherit data bias and can encode problematic perceptual distinctions; “Latent Compass” explicitly notes that the dimensions users explore reflect both their own biases and those of the GAN’s training data [2012.14283]. Robotics systems depend on known object models, accurate calibration, and computationally expensive simulation [2011.11397] [2509.15717]. Neuroscientific decoding of imagination raises questions of consent and mental privacy [2404.05468]. Taken together, these works suggest that imaginative perception is best understood not as a single faculty but as a class of constructive mechanisms—metaphorical, neural, generative, and agentic—by which absent structure becomes perceptually available.

Source: https://www.emergentmind.com/topics/imaginative-perception