Papers
Topics
Authors
Recent
Search
2000 character limit reached

ThematicPlane: Semantic Image Generation

Updated 8 July 2026
  • ThematicPlane is a system that allows users to navigate high-level semantic concepts such as mood and style to manipulate image generation.
  • It employs a multi-stage process including theme extraction, semantic axis construction, and diffusion-based synthesis to align user intent with model outputs.
  • User studies reveal an effective balance between exploratory and directed creative modes, while highlighting the need for enhanced explainability and multidimensional control.

ThematicPlane is a system for image generation that is designed to bridge users’ tacit, high-level creative intent and the control required to manipulate latent generative spaces in image editing (Lee et al., 8 Aug 2025). Rather than requiring ideas to be externalized through verbose prompts or reference images, it enables navigation of semantic concepts such as mood, style, or narrative tone within an interactive thematic design plane. The system is positioned as a semantics-driven interaction paradigm in which users iteratively steer outputs through thematic axes, reuse generated images as new references, and move between divergent and convergent creative modes. Its reported contribution is not only an interface technique but also an account of how such an interface mediates the gap between tacit intent and model behavior, including the need for more explainable controls when semantic mappings are not interpreted as expected (Lee et al., 8 Aug 2025).

1. Problem formulation and conceptual basis

ThematicPlane addresses a specific limitation in contemporary generative image tools: existing systems often require users to externalize ideas through prompts or references, which can constrain fluid exploration (Lee et al., 8 Aug 2025). In the formulation associated with the system, the central obstacle is the “externalization assumption,” namely the requirement that creative goals be articulated in explicit textual or parametric form before the system can act on them. ThematicPlane instead proposes that users should be able to operate directly on high-level semantic concepts.

The system’s central concept is “tacit-to-latent bridging,” defined as allowing users to navigate and manipulate high-level semantic concepts directly within an interactive thematic design plane (Lee et al., 8 Aug 2025). The operative semantic units are themes such as mood, style, or narrative tone. This shifts control from low-level prompt engineering to an interaction model in which semantic movement within an interface is the primary mechanism for producing variations in generated imagery.

The architecture is described as being grounded in a semantic interaction model inspired by the Circumplex Model of affect (Lee et al., 8 Aug 2025). Within that framing, the thematic plane functions as a multidimensional, interactive scaffold for editing. This suggests that ThematicPlane is not merely a prompt wrapper, but a structured semantic control surface intended to mediate between user intention and latent-space operations.

2. System architecture and semantic mapping pipeline

ThematicPlane’s technical pipeline begins with an input image and proceeds through theme extraction, semantic axis construction, embedding-based ranking, and image generation (Lee et al., 8 Aug 2025). The pipeline couples language-based semantic analysis with vision-language embedding alignment and diffusion-based image synthesis.

Stage Operation Components
Input and theme extraction Users input an image; descriptive keywords are extracted; object-based terms are filtered out GPT-4o
Semantic axis construction The system creates 12 thematic perturbations per theme, mapped along semantic axes Thematic plane
Embedding and mapping Theme descriptors are ranked by alignment with visual content DINOv2 and its aligned text encoder; cosine similarity
Image generation and iteration Top-ranked descriptors are injected into prompts; users can iterate from generated outputs Imagen 3

In the first stage, a LLM, GPT-4o, extracts descriptive keywords from the image, after which object-based terms are filtered out so that only thematic elements are retained (Lee et al., 8 Aug 2025). The retained descriptors are then used to construct semantic axes. The reported implementation creates 12 thematic perturbations per theme, organized along left/right semantic directions to form the plane for exploration.

The alignment step uses DINOv2 for images and its aligned text encoder for themes, with cosine similarity used to rank how closely each theme-related descriptor aligns with the visual content (Lee et al., 8 Aug 2025). The top-ranked thematic descriptors are then injected into image generation prompts using Imagen 3, described as a modern diffusion model. Iteration is built into the architecture: any generated image can be selected as a new reference, allowing stepwise semantic navigation rather than a single-shot generation process.

3. Interface model and modes of interaction

The principal interface component is the thematic interaction plane, which serves as the central visualization through which users navigate or steer across identified semantic axes (Lee et al., 8 Aug 2025). This design treats semantic movement as the core interaction primitive. The axes may correspond to shifts such as moving a mood from somber to joyful, although the broader claim is that the plane supports manipulation of thematic dimensions rather than a fixed vocabulary.

The interface also provides interactive visual feedback after each edit. Image generations reflecting the selected semantic direction are previewed, and users can iteratively select, compare, and further adjust outputs (Lee et al., 8 Aug 2025). This interaction loop is important to the system’s semantics-driven design: the plane is not simply descriptive metadata attached to an image, but an active control surface that mediates repeated cycles of generation and refinement.

The semantic controls are explicitly framed as alternatives to technical parameters. Users interact by choosing or dragging across conceptual dimensions rather than manipulating low-level model settings (Lee et al., 8 Aug 2025). The intended effect is alignment with conceptual ways of thinking about images. A plausible implication is that the interface attempts to preserve the user’s internal creative framing while minimizing the translation overhead typically imposed by prompt authoring.

4. Exploratory study and observed creative workflows

ThematicPlane was evaluated through an exploratory study with N=6N=6 participants aged 19–31, of mixed gender, with prior experience using generative tools (Lee et al., 8 Aug 2025). Participants used ThematicPlane alongside two baseline ChatGPT conditions to complete image transformation objectives. The study is therefore framed as comparative and exploratory rather than confirmatory.

A central reported finding is that participants moved between divergent and convergent creative modes (Lee et al., 8 Aug 2025). Divergent use involved open-ended exploration and the search for inspiration, whereas convergent use involved goal-driven refinement. The system supported both patterns within the same interaction framework. This duality is significant because it situates ThematicPlane between ideation support and directed editing rather than restricting it to one class of creative task.

The study also reports that users embraced unexpected results as inspiration or iteration cues (Lee et al., 8 Aug 2025). Even when outputs were not fully predictable, users incorporated them into their workflows. The reported predictability rating was M=3.43M=3.43, SD=1.27SD=1.27, while satisfaction remained high at M=5.86M=5.86, SD=0.90SD=0.90. These findings indicate that low predictability did not preclude positive evaluations, provided the system supported iterative reuse and semantic steering.

5. Explainability, interpretive variability, and control

A recurrent issue in the study is the variability in how participants understood themes and how they expected themes to map to outputs (Lee et al., 8 Aug 2025). Participants grounded their exploration in familiar themes, but they did not necessarily share the same assumptions about the semantics of those themes. Some expected a linear mapping from semantic changes to image outputs; others found the mapping less clear.

This variability led to an explicit call for greater explainability (Lee et al., 8 Aug 2025). The challenge is not only model transparency in the abstract, but the more specific problem of making the thematic plane’s semantics legible to users whose internal taxonomies of mood, style, or tone differ. In this sense, explainability is tied to alignment between user interpretation and system interpretation, not merely to disclosure of model internals.

The study also recorded a desire for greater control, including interest in multidimensional navigation beyond a single axis (Lee et al., 8 Aug 2025). That finding is technically important because the system is already described as a multidimensional scaffold, yet the user feedback indicates that the currently instantiated interaction did not fully satisfy all expectations for multidimensional semantic manipulation. This suggests a tension between semantic abstraction and controllability: higher-level controls may reduce prompt burden, but they also require sufficiently transparent mappings to remain intelligible in practice.

6. Significance within semantics-driven generative design

ThematicPlane is presented as a semantics-driven interaction system that seeks to make generative design more accessible, expressive, and iterative by enabling users to operate on their own conceptual terms (Lee et al., 8 Aug 2025). In that framing, its contribution lies less in a new generative backbone than in a new interface layer connecting tacit intent to latent-space traversal.

The reported findings support the view that creative work with generative systems is not exhausted by direct specification of goals. Participants used unexpected generations as opportunities for redirection and refinement, which positions serendipity as an integral component of the workflow rather than a failure mode (Lee et al., 8 Aug 2025). This suggests that ThematicPlane treats generation as a dialogic process in which semantic navigation and model output co-constitute the next design step.

At the same time, the study remains exploratory, and the system’s own findings foreground unresolved issues of explainability, semantic ambiguity, and multidimensional control (Lee et al., 8 Aug 2025). Those limitations are part of its significance. ThematicPlane highlights a broader research direction in which intuitive, semantics-driven interaction is treated as a primary design problem for generative tools, especially when aligning outputs with nuanced creative intent.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ThematicPlane.