---
title: 'Sensorium Arc: Multimodal Oceanic AI System'
url: https://www.emergentmind.com/topics/sensorium-arc
type: topic
---

# Sensorium Arc: Multimodal Oceanic AI System

Sensorium Arc is a real-time, multimodal interactive AI agent system designed for oceanic data exploration and interactive eco-art. Developed in partnership with the Center for the Study of the Force Majeure and influenced by Newton Harrison’s eco-aesthetic philosophy, Sensorium Arc operationalizes the personification of the ocean as a poetic, conversational speaker. The system enables users to engage in natural, spoken dialogues with an “Ocean” AI agent, which responds with a blend of scientific insight and ecological poetry, dynamically grounding content in high-dimensional environmental data. The architecture employs a modular, multi-agent approach, retrieval-augmented large language models (LLMs), formal grammar parsing, and real-time audiovisual rendering to facilitate immersive exploration and affective access to complex marine datasets [2511.15997].

## 1. System Architecture

Sensorium Arc is organized as a modular, multi-agent pipeline embedded in Unity, with four primary stages: Input Processing, LLM Pipeline, Response Processing, and Audio-Visual Layers. The pipeline operates as follows:

- **Input Processing:** 
  - Utilizes a proximity sensor (Arduino, $R = 0.5$ m threshold), Whisper-style microphone with active noise cancellation, and Whisper-Tiny speech-to-text (STT) for real-time audio capture and transcription.
- **LLM Pipeline:**
  - Composed of three stateless agents:
    1. **Visualization Decider Agent** determines which visual resources to trigger.
    2. **Query Rewriter + Retriever Agent** reformulates user queries and performs retrieval over the corpora.
    3. **Responder Agent** generates the final response as the “Ocean.”
- **Response Processing:**
  - Maps LLM-generated tokens to visualization events and prepares Unity-based text-to-speech (TTS) via Unity Jets.
- **Audio-Visual Layers:**
  - Includes globe visualizations (NASA EarthData), pre-rendered video overlays (e.g., plankton blooms, plastic dispersion), and synchronized subtitles.

All modules communicate asynchronously via JSON objects and are parallelized across GPU/CPU threads for real-time interaction.

## 2. Retrieval-Augmented LLM Integration

Sensorium Arc employs a retrieval-augmented, retrieve-then-generate LLM approach with $k=2$ nearest-neighbor document chunks. Data preprocessing involves segmenting the “Harrison Corpus”—a composite of manifestos and scientific papers—into sentences, embedding using all-MiniLM-L12-v2 ($\mathbb{R}^{384}$), and indexing via Approximate Nearest Neighbors (usearch/HNSW).

The real-time inference workflow is as follows:

1. **Query Rewriting:** Reformulate raw user query $q$ into “clean” query $q'$ using a chain-of-thought prompt in Qwen 8B.
2. **Embedding:** Embed $q'$ as $e_q \in \mathbb{R}^{384}$.
3. **Nearest Neighbor Retrieval:** Retrieve top-$k$ sentences $i_1$, $i_2$ by maximizing cosine similarity:
   $$
   \mathrm{sim}(e_q, e_i) = \frac{\langle e_q, e_i \rangle}{\|e_q\| \cdot \|e_i\|}
   $$
   with $\{i_1, i_2\} = \arg\max_{i \in DB} \mathrm{sim}(e_q, e_i)$.
4. **Context Assembly:** Retrieve full paragraphs $P_1$, $P_2$ containing $i_1$, $i_2$, which are prepended to the context window for the Responder Agent, grounding responses in sourced narrative.

## 3. Multimodal Data and Event Triggers

The data ingestion pipeline encompasses four major sources:

- **Time-series oceanographic measurements** (CO₂, chlorophyll, SST, currents, Kd)
- **Geospatial globe meshes** (latitude/longitude mapped textures)
- **Pre-rendered video layers** (plankton blooms, plastic dispersion, sea-level animation)
- **Text-to-speech (TTS) output**

Visualization and playback are triggered via two mechanisms:

- **Keyword Detection:** 
  - A fixed list of tokens (e.g., “chlorophyll,” “plastic,” “acidification”) mapped to visualization layers.
  - Pseudocode:
    ```python
    VISUAL_TOKENS = {"chlorophyll", "plastic", "acidification", ...}
    def parse_for_triggers(text):
        triggers = set()
        for token in tokenize(text):
            if token.lower() in VISUAL_TOKENS:
                triggers.add(token.lower())
        return triggers
    ```
- **Semantic Parsing:** 
  - GBNF grammar extracts temporal and regional phrases from user queries.
    ```bnf
    <Query>        ::= <TimeClause> <LocationClause> <Rest>
    <TimeClause>   ::= "in" <Year>
    <Year>         ::= /\d{4}/
    <LocationClause> ::= "at" <RegionName>
    <RegionName>     ::= "North Atlantic" | "Mediterranean" | ...
    ```
  - A recursive-descent parser identifies (time, location) pairs for dynamic texture and camera selection.

## 4. Conversational Design and Training

The “Ocean” persona is instantiated in the Responder Agent, which maintains the following per-turn context:
$$
\text{context} = \{ \text{user\_query},\ \text{vis\_token}, P_1, P_2, \text{conversation\_history} \}
$$
While intent is not explicitly modeled, chain-of-thought prompts direct the agent to answer scientific questions, weave ecological poetry, and reference paragraphs $P_1$, $P_2$.

The Responder is trained on standard cross-entropy loss,
$$
L = -\sum_{t=1}^T \log p(y_t \mid y_{<t},\ \text{context})
$$
with additional reward-shaping heuristics penalizing responses that exceed maximum length or omit any retrieved fact:
$$
\text{Reward} = -\alpha \cdot \text{LengthPenalty} + \beta \cdot \text{HasFact}(P_1) + \beta \cdot \text{HasFact}(P_2)
$$
Parameters $\alpha$, $\beta$ are tuned empirically in RL or via prompting.

Prompt templates support alternation between scientific statements and poetic lines, operationalizing the eco-poetic style:
```
“Ocean says: <SCIENTIFIC_STATEMENT>.
   And in my memory … <POETIC_LINE>.”
```

## 5. Real-Time Visualization and Audiovisual Playback

After response generation:

- **Layer Mapping:** Visualization Decider’s selected token (e.g., “chlorophyll”) activates corresponding globe or video layer via:
  ```
  VISUAL_MAP = {
      "chlorophyll" : CHLORO_LAYER,
      "plastic"     : PLASTIC_VIDEO,
      ...
  }
  ```
- **Camera Control:** Token triggers map to Cinemachine camera motions. E.g., on “chlorophyll” trigger: activate layer and set camera position (lat, lon, zoom).
- **Temporal Mapping:** Year cues map to frames using:
  $$
  \text{frame\_index} = \mathrm{round}\left( \frac{\mathrm{year} - 1990}{2024 - 1990} \cdot (N_{\text{frames}} - 1) \right)
  $$
- **Audio-Visual Synchronization:** Master clock coordinates TTS audio and sentence-level subtitles, chunked at sentence boundaries.

## 6. Eco-Aesthetic and Narrative Approach

Sensorium Arc embraces Newton Harrison’s eco-aesthetic philosophy, reconceptualizing ocean data as narrative symbols rather than mere quantitative abstractions. Key operationalizations include:

- Equal weighting of Harrisons’ manifestos and scientific literature during retrieval.
- Chained prompt templates alternating data-driven commentary with ekphrastic, poetic imagery (e.g., “I carry sunlight in my plankton blooms”).
- Respect for layered temporality, integrating notions of past and future (Force Majeure) within a single narrative.

This paradigm supports the mediation of intuitive, affective access to high-dimensional marine datasets and foregrounds the personification of the ocean as a narrative entity.

## 7. Prototypical Interaction Workflow

A typical user-agent exchange proceeds as follows:

| Stage                      | Description                                                          | Mechanism/Output                                                                       |
|----------------------------|-----------------------------------------------------------------------|-----------------------------------------------------------------------------------------|
| Input Processing           | STT transcription on proximity sensor trigger                         | $q = $ “Ocean, what happened to plastic pollution in the North Atlantic around 2010?”    |
| Visualization Decider      | Few-shot prompted classification                                      | Outputs `VISUAL: plastic`                                                               |
| Query Rewriter             | Chain-of-thought reasoning, marker stripping                         | $q' = $ “plastic pollution North Atlantic 2010 trend”                                   |
| Retrieval                  | Embedding and nearest-neighbor search                                | Retrieves $P_1$, $P_2$ on plastic dispersion                                            |
| Responder                  | LLM generation with context, scientific and poetic blending           | Generates Ocean’s answer, e.g. “7% per year” plastics increase, poetic current metaphor  |
| Response Processing        | Maps trigger to visualization and extracts year cue                   | Triggers PLASTIC_VIDEO, sets globe frame to 2010, TTS output                            |
| Audio-Visual Layers        | Visualization/camera activation, subtitles, synchronized playback     | Dynamic globe video, camera pan, captions in sync with TTS                              |

The result is a multimodal experience in which the Ocean agent responds with grounded data and metaphor while the visualization layer projects scientifically accurate, time-specific environmental changes.

---

Sensorium Arc exemplifies the integration of modular LLM-agent architectures, retrieval-enhanced groundedness, formal grammar pipelines, and dynamic Unity-based rendering for immersive, ecopoetic interaction with oceanic data [2511.15997].

Source: https://www.emergentmind.com/topics/sensorium-arc