Papers
Topics
Authors
Recent
Search
2000 character limit reached

ASAL++: Evolving Targets in Artificial Life

Updated 12 July 2026
  • ASAL++ is a multimodal framework that evolves target prompts using a two-model system to enable open-ended search in artificial life simulations.
  • It employs two strategies—Evolved Supervised Targets (EST) for high visual novelty and Evolved Temporal Targets (ETT) for coherent evolutionary narratives, with measured OE improvements.
  • Experiments in Lenia demonstrate that evolving prompts dynamically over iterations produces both visually novel and interpretable evolutionary trajectories.

ASAL++ most explicitly denotes “Automated Search for Artificial Life PlusPlus”, a method for open-ended-like search in artificial life simulations guided by multimodal foundation models. It extends ASAL by adding a second foundation model that proposes new evolutionary targets from a simulation’s visual history, thereby inducing an evolutionary trajectory with increasingly complex targets. The framework was introduced for Lenia and compared through two target-evolution strategies, Evolved Supervised Targets (EST) and Evolved Temporal Targets (ETT), with EST promoting greater visual novelty and ETT fostering more coherent and interpretable evolutionary sequences (Baid et al., 26 Sep 2025). In a separate literature on VR video quality assessment, ASAL++ also names the continual-learning extension of Adaptive Score Alignment Learning, indicating that the designation is context-dependent across domains (Zhou et al., 27 Feb 2025).

1. Definition and lineage

ASAL++ builds directly on ASAL, a substrate-agnostic framework that uses vision-language foundation models to automatically explore large spaces of artificial life simulations. In ASAL, a substrate SS is treated as a parametrized family of simulations, and a vision-LLM aligns rendered simulation outputs with natural-language prompts or evaluates novelty and diversity in embedding space. The original ASAL formulation is organized around three search modes: finding simulations that produce target phenomena, discovering simulations that generate temporally open-ended novelty, and illuminating an entire space of interestingly diverse simulations (Kumar et al., 2024).

In this lineage, ASAL++ retains the ASAL inner loop—CLIP-based evaluation combined with evolutionary optimization over simulation parameters—but alters the status of the target itself. Rather than fixing the prompt sequence in advance, ASAL++ lets a second multimodal foundation model propose the next target prompt after observing the current optimized simulation. The outer loop therefore modifies the optimization objective over time, while the inner loop continues to optimize the substrate parameters against the current prompt or prompt sequence. This is the defining step that moves the method from directed search toward open-ended-like search (Baid et al., 26 Sep 2025).

A central motivation is that ASAL’s prompts are pre-specified and static: the targets are chosen upfront by a human, the optimization only searches for simulations that match these fixed targets, and there is no automatic mechanism for discovering intermediate stepping stones or progressively more complex objectives. ASAL++ addresses this by making the target prompts themselves evolve over time and by letting a second foundation model propose new targets based on the visual history of simulations (Baid et al., 26 Sep 2025).

2. Formal architecture and objective structure

ASAL++ is a two-FM system. The first foundation model is a vision-LLM (CLIP), which acts as an evaluator by mapping simulation frames and text prompts into a shared latent space. The second is an EvolverModel—in the reported experiments, Gemma-3-4b-it—which accepts the rollout video of the current best simulation and produces a new natural-language target prompt (Baid et al., 26 Sep 2025).

The substrate formalism follows ASAL. A simulation is determined by parameters θ\theta and consists of an initial state distribution, forward dynamics, and a rendering function. Running the simulation for TT steps defines a rollout

RST(θ)=Renderθ(StepθT(s0)).RS^{T}(\theta) = \text{Render}_\theta\big(\text{Step}_\theta^{T}(s_0)\big).

ASAL++ often works with the entire image sequence

Imgs0:T=(Renderθ(s0),Renderθ(s1),…,Renderθ(sT)).\text{Imgs}_{0:T} = \big(\text{Render}_\theta(s_0), \text{Render}_\theta(s_1), \dots, \text{Render}_\theta(s_T)\big).

If E\mathcal{E} denotes the encoder derived from the VLM, the inner-loop loss for a single prompt is

L(θ)=−⟨z0:T,zp⟩,\mathcal{L}(\theta) = - \left\langle z_{0:T}, z_p \right\rangle,

where z0:T=E(Imgs0:T)z_{0:T} = \mathcal{E}(\text{Imgs}_{0:T}) and zp=E(p)z_p = \mathcal{E}(p). Maximizing image–text similarity is therefore equivalent to minimizing L(θ)\mathcal{L}(\theta) (Baid et al., 26 Sep 2025).

To quantify within-simulation visual novelty, the method defines an OE score

θ\theta0

This compares the final frame at time θ\theta1 to all earlier frames in CLIP space, takes the maximum similarity, and subtracts it from θ\theta2. Higher OE therefore means the final frame is more visually novel relative to its own history (Baid et al., 26 Sep 2025).

The resulting architecture is a nested loop. The inner loop is ASAL: CLIP provides a similarity-based fitness, and Sep-CMA-ES optimizes θ\theta3. The outer loop is prompt evolution: Gemma-3 observes the optimized rollout and proposes the next target prompt. The best parameters θ\theta4 from the current outer iteration are then used as a warm start for the next iteration (Baid et al., 26 Sep 2025).

3. Evolved targets: EST and ETT

ASAL++ introduces two strategies for target evolution. In Evolved Supervised Targets (EST), the new prompt proposed at iteration θ\theta5 replaces the previous target. The next inner loop therefore optimizes the simulation only for that single new prompt. Past prompts influence the process only through the inherited parameters and the visual content that Gemma-3 sees (Baid et al., 26 Sep 2025).

In Evolved Temporal Targets (ETT), the new prompt is appended to the list of all previous prompts. The simulation is then optimized to match the entire prompt sequence temporally. The prompts become a narrative of evolutionary stages, and Gemma-3 conditions on both the rollout video and the full prompt history when proposing the next target (Baid et al., 26 Sep 2025).

The operational difference is substantial. EST effectively jumps to a new single objective at each outer iteration. This encourages high novelty between iterations, but it weakens long-term coherence. ETT instead forces later simulations to satisfy all prompts in the growing sequence, so earlier objectives cannot simply be discarded. The intended effect is greater continuity, greater interpretability, and a more legible evolutionary story (Baid et al., 26 Sep 2025).

The target-generation interface is correspondingly explicit. In the ETT setting, the instruction to the EvolverModel includes the sequence of prompts used so far, treats them as “ecological niches that have already been explored,” specifies the current iteration number, and asks for the “NEXT TARGET PROMPT” while encouraging a direction that is “significantly different from the past” and “interesting lifelike behaviour.” The model is instructed to output only the new prompt string and to remain concise (Baid et al., 26 Sep 2025).

4. Experimental realization in Lenia

The reported experiments use Lenia as the substrate. Lenia is described as a continuous generalisation of Conway’s Game of Life with continuous state values, continuous time and space, a radially symmetric growth kernel, and a rich diversity of self-organising, lifelike patterns such as gliders, oscillators, and orbiters. It was chosen for its expressivity, efficiency, and prior importance within artificial life research (Baid et al., 26 Sep 2025).

The experimental configuration is specific. The rollout length is θ\theta6 time steps, so each evolved simulation consists of 256 frames. The number of ASAL++ outer iterations is θ\theta7, and the inner ASAL optimization uses θ\theta8 CMA-ES steps per outer iteration. The initial human prompts are nine phrases: “a pepperoni pizza,” “a slime mould,” “a flower,” “the garden of eden,” “an extraterrestrial life,” “a monkey,” “a caterpillar,” “a microbe,” and “a fungus” (Baid et al., 26 Sep 2025).

The optimization stack is also fixed in the reported system. The evaluator is CLIP, used both for prompt alignment and for OE scoring. The prompt generator is Gemma-3-4b-it, used in multimodal mode with the full rollout video from the previous iteration. In the “trees of life” experiments, GPT-4 is additionally used to generate environmental prompt pairs such as “high energy” and “low energy” (Baid et al., 26 Sep 2025).

At the algorithmic level, ASAL++ initializes CMA-ES with population size θ\theta9 and step size TT0, embeds the initial prompt via TT1, and then alternates between inner-loop optimization and outer-loop prompt generation. For ETT, the prompt update is TT2; for EST, the update is TT3. In both cases, the next inner loop starts from the current best simulation parameters TT4 (Baid et al., 26 Sep 2025).

5. Quantitative behavior, qualitative trajectories, and “trees of life”

The reported metrics distinguish the two target-evolution strategies clearly. For EST, the mean OE across initial prompts is TT5, and the mean TT6 is TT7. For the initial prompt “an extraterrestrial life”, the paper reports TT8 and TT9. These results are associated with larger jumps in morphology between iterations and stronger exploration of visually distinct niches (Baid et al., 26 Sep 2025).

For ETT, the mean OE is RST(θ)=Renderθ(StepθT(s0)).RS^{T}(\theta) = \text{Render}_\theta\big(\text{Step}_\theta^{T}(s_0)\big).0 and the mean RST(θ)=Renderθ(StepθT(s0)).RS^{T}(\theta) = \text{Render}_\theta\big(\text{Step}_\theta^{T}(s_0)\big).1 is RST(θ)=Renderθ(StepθT(s0)).RS^{T}(\theta) = \text{Render}_\theta\big(\text{Step}_\theta^{T}(s_0)\big).2. The reported qualitative interpretation is different from EST: OE remains roughly similar to baseline ASAL within error, but the resulting sequences exhibit clearer continuity. A representative example begins from “a microbe” and evolves through prompts such as “clusters of microbe-like entities” and “microbe motility in fluid-like environment,” producing an interpretable sequence of formation, clustering, and coordinated motion (Baid et al., 26 Sep 2025).

The paper also describes “trees of life”, where Gemma-3 is sampled multiple times at higher temperature to branch the prompt sequence. Starting from “a caterpillar,” branches diverge rapidly, with some preserving elongated segmented patterns and others evolving radial or fractal-like structures. When environmental descriptors generated by GPT-4 are added—examples include “high energy,” “low energy,” “expansive,” “conservative,” “stable,” and “unstable”—different branches produce visually distinct simulations, with “high energy” branches showing turbulent, more fragmented patterns and “low energy” branches showing calmer, more homogeneous patterns (Baid et al., 26 Sep 2025).

These results are presented as evidence that ASAL++ injects open-ended-like characteristics into artificial life search. The fitness function is no longer fixed, each new prompt functions as a stepping stone, and the EvolverModel acts as a surrogate for a human scientist that observes the current system and proposes a new direction. EST emphasizes continual novelty, whereas ETT emphasizes coherent evolutionary narratives (Baid et al., 26 Sep 2025).

6. Interpretation, limitations, and terminological scope

The authors state explicitly that ASAL++ does not guarantee truly unbounded open-ended evolution. The open-endedness claim is empirical and is limited by the substrate, model biases, and computational horizons. Several concrete limitations are identified: Gemma-3 often ignores the literal initial prompt when it is non-biological; Lenia’s radial symmetry and Gaussian growth favour circular patterns; each outer iteration requires 2000 inner optimization steps with CMA-ES; ETT sometimes converges to local minima and can produce prompt repetition; and CLIP processes single images, so low-resolution frames may miss temporal dynamics. Suggested directions include more expressive substrates, larger or fine-tuned multimodal foundation models, video-specific VLMs, quality-diversity algorithms such as MAP-Elites or AURORA, dynamic stopping via OE, multi-agent interactions, and human-in-the-loop steering (Baid et al., 26 Sep 2025).

The designation ASAL++ is not confined to this ALife setting. In a distinct VR-VQA literature, it denotes the continual-learning extension of Adaptive Score Alignment Learning, combining correlation loss with error loss, a Gaussian score layer, key frame extraction, feature adaptation, and adaptive memory replay under computation and storage restrictions of VR devices (Zhou et al., 27 Feb 2025). The designation also appears as an informal label for Anderson-accelerated augmented Lagrangian / ADMM schemes in extended waveform inversion, where an augmented Lagrangian fixed-point iteration is accelerated by Anderson acceleration (Aghazade et al., 2021).

This suggests that ASAL++ is presently a context-dependent label rather than a single cross-domain formalism. Within the literature that explicitly introduces the term in its title, however, ASAL++ denotes a multimodal, two-foundation-model framework for evolving objectives in artificial life search: CLIP evaluates simulated rollouts against prompts, Gemma-3 proposes the next prompt from visual history, and the resulting outer–inner loop induces either novelty-oriented or coherence-oriented evolutionary trajectories in Lenia (Baid et al., 26 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ASAL++.