---
title: 'Wandering Prototypes: Adaptive Online Memory'
url: https://www.emergentmind.com/topics/wandering-prototypes
type: topic
---

# Wandering Prototypes: Adaptive Online Memory

Wandering prototypes are dynamic class representations in feature space, induced by continual interaction with temporally evolving, context-rich data streams. The term emerges in the context of online contextualized few-shot learning, where prototypes—class summary vectors—undergo continuous adaptation to shifts in spatiotemporal context, enabling robust classification and novelty detection as agents “wander” through changing environments [2007.04546]. This adaptive mechanism is formalized in architectures such as the Contextual Prototypical Memory (CPM) model, where prototype “wandering” denotes the incremental, context-driven drift of memory representations in the embedding space.

## 1. Foundations of Online Contextualized Few-Shot Learning

The paradigm extends classical few-shot learning from episodic, stationary evaluation into a continual, temporally structured domain. At each time step $t=1,…,T$, the learner receives an input $x_t\in\mathbb{R}^d$ (e.g., an image feature vector), a possibly missing class label $\tilde{y}_t\in\{1,…,K_t,-1\}$ (where $K_t$ is the number of known classes and $-1$ encodes unlabeled instances), and is implicitly associated with an unobserved context index $c_t\in\{1,…,C\}$. The model must simultaneously:

- Classify known vs. novel classes: Predict whether $x_t$ belongs to a previously seen class or a new class.
- Identify class label for known instances.
- Integrate new classes as they appear, without a fixed upper bound on $K_t$.

Critically, the context $c_t$—such as the agent’s environment or “room”—evolves according to a Markov process, rendering class distributions and feature structures non-stationary. There is no distinction between train and test phases; every step requires simultaneous adaptation and evaluation [2007.04546].

## 2. Latent Context Inference via Recurrent Encoding

Context classification is not provided as direct supervision but is inferred from feature streams. Context transitions evolve as a Markov chain:
$$
p(c_t | c_{t-1}) = (1-p_s) \cdot 1[c_t = c_{t-1}] + p_s \cdot (1/(C-1)) \cdot 1[c_t \neq c_{t-1}]
$$
where $p_s$ is the switching probability. Observations $x_t$ provide weak evidence for current context; proper context inference depends on sequence modeling.

The CPM model operationalizes context summarization as follows:
- Features $h_t^{\text{cnn}}$ are extracted via a convolutional neural network.
- A recurrent neural network (RNN) encoder produces a hidden state $z_t$ and a context embedding $h_t^{\text{rnn}}$.
- The context summary vector $h_t = h_t^{\text{cnn}} + h_t^{\text{rnn}}$ fuses perceptual and temporal/context cues, modulating all downstream memory operations [2007.04546].

## 3. Construction and Evolution of Contextual Prototypes

Each observed class-context pair $(k, c)$ is represented by a prototype vector $p_{k,(c)} \in \mathbb{R}^d$ and an update count $n_{k,(c)}$. The prototype update for a new observation $(x_t, y_t=k, c_t)$ takes the form:
$$
n_{k,(c_t)} \leftarrow n_{k,(c_t)} + 1
$$
$$
p_{k,(c_t)} \leftarrow \frac{(n_{k,(c_t)} - 1) \cdot p_{k,(c_t)} + h_t}{n_{k,(c_t)}}
$$
where $A(h; p, n) = (n p + h)/(n + 1)$ denotes the online averaging operation.

In practice, CPM does not explicitly maintain a context-indexed prototype table, but modulates all prototypes by the context embedding $h_t^{\text{rnn}}$, causing class prototypes $p_k$ to drift in feature space as the RNN’s representation of context evolves. The prototype’s “wandering” thus encodes both cumulative evidence and ongoing environmental shifts.

## 4. Prototype Retrieval, Novelty Detection, and Online Update

At each step, the full procedure involves context-sensitive retrieval and updating:

- *Embedding:* The agent encodes $x_t$ to $h_t$ using the context-sensitive RNN-CNN fusion.
- *Similarity computation:* Scaled distances $d_{k}= \| (m_t \odot h_t) - p_k\|^2$ are computed for each prototype, where $m_t$ is a metric scaling factor.
- *Novelty detection:* Predict old vs. new class via
  $$
  \hat{u}_t^r = \sigma\left( \frac{\min_k d_k - \beta_t^r}{\gamma_t^r} \right)
  $$
  with thresholding. Here, $\beta_t^r, \gamma_t^r$ are RNN-controlled read parameters.
- *Classification:* Softmax assignment among old classes: $\hat{y}_{t,[k]} \propto \exp(-d_k)$.
- *Controlled prototype update:* Write strengths for updating or creating prototypes are determined by RNN control parameters and labeling status.
- *Online update:* The moving average is applied; if novelty confidence is high, a new prototype is created.

This mechanism ensures that prototypes gradually reflect both recently observed features and their fluctuating spatiotemporal context, supporting rapid recognition without catastrophic forgetting [2007.04546].

## 5. Datasets and Empirical Observations of Prototype Wandering

Dedicated benchmarks simulate sequential, context-dependent recognition problems:

| Benchmark         | # Classes | Contexts/Episode | Description                                 |
|-------------------|-----------|------------------|---------------------------------------------|
| RoamingOmniglot   | 6492      | 5–10             | Handwritten characters, 150-frame episodes  |
| RoamingImageNet   | 608       | —                | Tiered-ImageNet subset, similar sampler     |
| RoamingRooms      | >7000     | 90               | 1.2M frames, robot navigation in 90 rooms   |

One-shot average precision (AP) (length 100–150) for CPM and baselines:

| Dataset            | CPM   | Online ProtoNet | DNC    |
|--------------------|-------|----------------|--------|
| RoamingOmniglot    | 94.2% | 90.5%          | 81.3%  |
| RoamingRooms       | 89.1% | 86.0%          | 80.9%  |
| RoamingImageNet    | 34.4% | 23.1%          | 26.8%  |

Empirically, as agents traverse environments, CPM’s context-modulated prototypes “wander” smoothly in embedding space. This wandering minimizes forgetting and supports stable discrimination under strongly non-stationary input streams, outperforming static and context-agnostic baselines by up to 5–8 AP points in various supervised and semi-supervised regimes [2007.04546].

## 6. Theoretical Connections and Broader Context

Wandering prototypes formalize a memory update mechanism that integrates new evidence while maintaining alignment with latent, unobserved contextual factors. Prototypes “wander” in response to the RNN’s context summary, such that memory representations track the evolving spatiotemporal state of the environment without explicit context segmentation.

This concept is related to dynamic representation learning, meta-learning with task inference, and continual learning under context drift. Unlike methods that freeze or rapidly overwrite class representatives, CPM’s framework enables prototypes to undergo continuous, context-informed plasticity, aligning nearest neighbor structure with the current operating context. A plausible implication is that this mechanism is particularly robust in settings where previously discriminative features become obsolete due to context transitions, yet catastrophic forgetting would hamper naive online adaptation [2007.04546].

## 7. Summary and Significance

Wandering prototypes embody an architecture for online classification and novelty detection in nonstationary, context-rich environments. Through a context-sensitive fusion of temporal and perceptual features, qualified by an RNN encoder, prototype vectors incrementally adapt, or “wander,” to reflect the most recent evidence and contextual priors. Empirical results indicate reduced forgetting and improved adaptation relative to context-agnostic and less plastic prototypical schemes. CPM thus demonstrates the importance of context-driven representational drift for online continual learning, formalizing a framework where wandering prototypes dynamically track a world in flux [2007.04546].

Source: https://www.emergentmind.com/topics/wandering-prototypes