---
title: 'ASAL++: Evolving Targets in Artificial Life'
url: https://www.emergentmind.com/topics/asal
type: topic
---

# ASAL++: Evolving Targets in Artificial Life

ASAL++ most explicitly denotes **“Automated Search for Artificial Life PlusPlus”**, a method for **open-ended-like search** in artificial life simulations guided by multimodal foundation models. It extends ASAL by adding a second foundation model that proposes new evolutionary targets from a simulation’s visual history, thereby inducing an evolutionary trajectory with increasingly complex targets. The framework was introduced for Lenia and compared through two target-evolution strategies, **Evolved Supervised Targets (EST)** and **Evolved Temporal Targets (ETT)**, with EST promoting greater visual novelty and ETT fostering more coherent and interpretable evolutionary sequences [2509.22447]. In a separate literature on VR video quality assessment, **ASAL++** also names the continual-learning extension of **Adaptive Score Alignment Learning**, indicating that the designation is context-dependent across domains [2502.19644].

## 1. Definition and lineage

ASAL++ builds directly on **ASAL**, a substrate-agnostic framework that uses vision-language foundation models to automatically explore large spaces of artificial life simulations. In ASAL, a substrate \(S\) is treated as a parametrized family of simulations, and a vision-language model aligns rendered simulation outputs with natural-language prompts or evaluates novelty and diversity in embedding space. The original ASAL formulation is organized around three search modes: finding simulations that produce target phenomena, discovering simulations that generate temporally open-ended novelty, and illuminating an entire space of interestingly diverse simulations [2412.17799].

In this lineage, ASAL++ retains the ASAL inner loop—CLIP-based evaluation combined with evolutionary optimization over simulation parameters—but alters the status of the target itself. Rather than fixing the prompt sequence in advance, ASAL++ lets a second multimodal foundation model propose the next target prompt after observing the current optimized simulation. The outer loop therefore modifies the optimization objective over time, while the inner loop continues to optimize the substrate parameters against the current prompt or prompt sequence. This is the defining step that moves the method from directed search toward **open-ended-like** search [2509.22447].

A central motivation is that ASAL’s prompts are pre-specified and static: the targets are chosen upfront by a human, the optimization only searches for simulations that match these fixed targets, and there is no automatic mechanism for discovering intermediate stepping stones or progressively more complex objectives. ASAL++ addresses this by making the target prompts themselves evolve over time and by letting a second foundation model propose new targets based on the visual history of simulations [2509.22447].

## 2. Formal architecture and objective structure

ASAL++ is a **two-FM system**. The first foundation model is a **vision-language model (CLIP)**, which acts as an evaluator by mapping simulation frames and text prompts into a shared latent space. The second is an **EvolverModel**—in the reported experiments, **Gemma-3-4b-it**—which accepts the rollout video of the current best simulation and produces a new natural-language target prompt [2509.22447].

The substrate formalism follows ASAL. A simulation is determined by parameters \(\theta\) and consists of an initial state distribution, forward dynamics, and a rendering function. Running the simulation for \(T\) steps defines a rollout
\[
RS^{T}(\theta) = \text{Render}_\theta\big(\text{Step}_\theta^{T}(s_0)\big).
\]
ASAL++ often works with the entire image sequence
\[
\text{Imgs}_{0:T} = \big(\text{Render}_\theta(s_0), \text{Render}_\theta(s_1), \dots, \text{Render}_\theta(s_T)\big).
\]
If \(\mathcal{E}\) denotes the encoder derived from the VLM, the inner-loop loss for a single prompt is
\[
\mathcal{L}(\theta) = - \left\langle z_{0:T}, z_p \right\rangle,
\]
where \(z_{0:T} = \mathcal{E}(\text{Imgs}_{0:T})\) and \(z_p = \mathcal{E}(p)\). Maximizing image–text similarity is therefore equivalent to minimizing \(\mathcal{L}(\theta)\) [2509.22447].

To quantify within-simulation visual novelty, the method defines an **OE score**
\[
\text{OE}(T; \theta) = 1 - \max_{T' < T} \left\langle \text{VLM}_{\text{img}\big(RS^{T}(\theta)\big)},\  \text{VLM}_{\text{img}\big(RS^{T'}(\theta)\big)} \right\rangle.
\]
This compares the final frame at time \(T\) to all earlier frames in CLIP space, takes the maximum similarity, and subtracts it from \(1\). Higher OE therefore means the final frame is more visually novel relative to its own history [2509.22447].

The resulting architecture is a nested loop. The **inner loop** is ASAL: CLIP provides a similarity-based fitness, and **Sep-CMA-ES** optimizes \(\theta\). The **outer loop** is prompt evolution: Gemma-3 observes the optimized rollout and proposes the next target prompt. The best parameters \(\theta^*\) from the current outer iteration are then used as a **warm start** for the next iteration [2509.22447].

## 3. Evolved targets: EST and ETT

ASAL++ introduces two strategies for target evolution. In **Evolved Supervised Targets (EST)**, the new prompt proposed at iteration \(n\) replaces the previous target. The next inner loop therefore optimizes the simulation only for that single new prompt. Past prompts influence the process only through the inherited parameters and the visual content that Gemma-3 sees [2509.22447].

In **Evolved Temporal Targets (ETT)**, the new prompt is appended to the list of all previous prompts. The simulation is then optimized to match the entire prompt sequence temporally. The prompts become a narrative of evolutionary stages, and Gemma-3 conditions on both the rollout video and the full prompt history when proposing the next target [2509.22447].

The operational difference is substantial. EST effectively jumps to a new single objective at each outer iteration. This encourages high novelty between iterations, but it weakens long-term coherence. ETT instead forces later simulations to satisfy all prompts in the growing sequence, so earlier objectives cannot simply be discarded. The intended effect is greater continuity, greater interpretability, and a more legible evolutionary story [2509.22447].

The target-generation interface is correspondingly explicit. In the ETT setting, the instruction to the EvolverModel includes the sequence of prompts used so far, treats them as “ecological niches that have already been explored,” specifies the current iteration number, and asks for the “NEXT TARGET PROMPT” while encouraging a direction that is “significantly different from the past” and “interesting lifelike behaviour.” The model is instructed to output only the new prompt string and to remain concise [2509.22447].

## 4. Experimental realization in Lenia

The reported experiments use **Lenia** as the substrate. Lenia is described as a continuous generalisation of Conway’s Game of Life with continuous state values, continuous time and space, a radially symmetric growth kernel, and a rich diversity of self-organising, lifelike patterns such as gliders, oscillators, and orbiters. It was chosen for its expressivity, efficiency, and prior importance within artificial life research [2509.22447].

The experimental configuration is specific. The rollout length is **\(T = 256\)** time steps, so each evolved simulation consists of **256 frames**. The number of ASAL++ outer iterations is **\(N = 8\)**, and the inner ASAL optimization uses **\(I = 2000\)** CMA-ES steps per outer iteration. The initial human prompts are nine phrases: “a pepperoni pizza,” “a slime mould,” “a flower,” “the garden of eden,” “an extraterrestrial life,” “a monkey,” “a caterpillar,” “a microbe,” and “a fungus” [2509.22447].

The optimization stack is also fixed in the reported system. The evaluator is **CLIP**, used both for prompt alignment and for OE scoring. The prompt generator is **Gemma-3-4b-it**, used in multimodal mode with the full rollout video from the previous iteration. In the “trees of life” experiments, **GPT-4** is additionally used to generate environmental prompt pairs such as “high energy” and “low energy” [2509.22447].

At the algorithmic level, ASAL++ initializes CMA-ES with population size \(P\) and step size \(\sigma\), embeds the initial prompt via \(\mathcal{E}\), and then alternates between inner-loop optimization and outer-loop prompt generation. For ETT, the prompt update is \(p \gets p + p'\); for EST, the update is \(p \gets p'\). In both cases, the next inner loop starts from the current best simulation parameters \(\theta^*\) [2509.22447].

## 5. Quantitative behavior, qualitative trajectories, and “trees of life”

The reported metrics distinguish the two target-evolution strategies clearly. For **EST**, the mean OE across initial prompts is **\(0.052 \pm 0.002\)**, and the mean \(\Delta\text{OE}\) is **\(0.009 \pm 0.004\)**. For the initial prompt **“an extraterrestrial life”**, the paper reports **\(\text{OE} \approx 0.0575 \pm 0.010\)** and **\(\Delta\text{OE} \approx 0.027\)**. These results are associated with larger jumps in morphology between iterations and stronger exploration of visually distinct niches [2509.22447].

For **ETT**, the mean OE is **\(0.049 \pm 0.002\)** and the mean \(\Delta\text{OE}\) is **\(0.002 \pm 0.004\)**. The reported qualitative interpretation is different from EST: OE remains roughly similar to baseline ASAL within error, but the resulting sequences exhibit clearer continuity. A representative example begins from “a microbe” and evolves through prompts such as “clusters of microbe-like entities” and “microbe motility in fluid-like environment,” producing an interpretable sequence of formation, clustering, and coordinated motion [2509.22447].

The paper also describes **“trees of life”**, where Gemma-3 is sampled multiple times at higher temperature to branch the prompt sequence. Starting from “a caterpillar,” branches diverge rapidly, with some preserving elongated segmented patterns and others evolving radial or fractal-like structures. When environmental descriptors generated by GPT-4 are added—examples include “high energy,” “low energy,” “expansive,” “conservative,” “stable,” and “unstable”—different branches produce visually distinct simulations, with “high energy” branches showing turbulent, more fragmented patterns and “low energy” branches showing calmer, more homogeneous patterns [2509.22447].

These results are presented as evidence that ASAL++ injects **open-ended-like characteristics** into artificial life search. The fitness function is no longer fixed, each new prompt functions as a stepping stone, and the EvolverModel acts as a surrogate for a human scientist that observes the current system and proposes a new direction. EST emphasizes continual novelty, whereas ETT emphasizes coherent evolutionary narratives [2509.22447].

## 6. Interpretation, limitations, and terminological scope

The authors state explicitly that ASAL++ does **not guarantee truly unbounded open-ended evolution**. The open-endedness claim is empirical and is limited by the substrate, model biases, and computational horizons. Several concrete limitations are identified: **Gemma-3 often ignores the literal initial prompt when it is non-biological**; **Lenia’s radial symmetry and Gaussian growth favour circular patterns**; each outer iteration requires **2000 inner optimization steps with CMA-ES**; **ETT sometimes converges to local minima** and can produce **prompt repetition**; and **CLIP processes single images**, so low-resolution frames may miss temporal dynamics. Suggested directions include more expressive substrates, larger or fine-tuned multimodal foundation models, video-specific VLMs, quality-diversity algorithms such as **MAP-Elites** or **AURORA**, dynamic stopping via OE, multi-agent interactions, and human-in-the-loop steering [2509.22447].

The designation **ASAL++** is not confined to this ALife setting. In a distinct VR-VQA literature, it denotes the continual-learning extension of **Adaptive Score Alignment Learning**, combining correlation loss with error loss, a Gaussian score layer, key frame extraction, feature adaptation, and adaptive memory replay under computation and storage restrictions of VR devices [2502.19644]. The designation also appears as an informal label for **Anderson-accelerated augmented Lagrangian / ADMM** schemes in extended waveform inversion, where an augmented Lagrangian fixed-point iteration is accelerated by Anderson acceleration [2106.14065].

This suggests that **ASAL++** is presently a context-dependent label rather than a single cross-domain formalism. Within the literature that explicitly introduces the term in its title, however, ASAL++ denotes a multimodal, two-foundation-model framework for evolving objectives in artificial life search: CLIP evaluates simulated rollouts against prompts, Gemma-3 proposes the next prompt from visual history, and the resulting outer–inner loop induces either novelty-oriented or coherence-oriented evolutionary trajectories in Lenia [2509.22447].

Source: https://www.emergentmind.com/topics/asal