---
title: Adaptive Relation Library (ARL)
url: https://www.emergentmind.com/topics/adaptive-relation-library-arl
type: topic
---

# Adaptive Relation Library (ARL)

Searching arXiv for the specified paper and closely related context.
Adaptive Relation Library (ARL) is a component of AutoLayout that functions as a collection of symbolic topological predicates together with numerical constraint functions and Boolean validation checks for layout synthesis. In this formulation, ARL describes how objects should be positioned or aligned, but differs from static, hand-designed rule sets because its entries are generated, parameterized, and continually refined by a large language model during the layout process. Within AutoLayout, ARL is part of a closed-loop self-validation mechanism embedded in a slow-fast collaborative reasoning framework, where it supports the generation and evaluation of layouts while mitigating spatial hallucination and balancing physical stability with semantic consistency [2507.04293].

## 1. Definition and motivation

In the context of layout synthesis, ARL is defined as a collection of symbolic topological predicates, their numerical constraint functions, and Boolean validation checks. The predicates encode positional and alignment relations such as anchoring, pairwise ordering, stacking, and multi-object alignment. The associated functions evaluate whether a candidate layout satisfies these relations either softly, through a score in $[0,1]$, or discretely, through a true/false validation [2507.04293].

The motivation for ARL is explicitly framed against static, hand-designed rule sets such as “object A must be left of B with at least 5 cm clearance.” ARL entries are instead generated, parameterized, and continually refined by a large language model during the layout process. Three advantages are identified. First, ARL provides flexibility across novel scenes because relations can be added or adjusted on-the-fly rather than hard-coded in advance. Second, it enables code-level self-repair: if a dynamically created constraint is too loose or too strict, the large language model can regenerate its Python implementation at runtime. Third, it supports a better semantic-physical trade-off by learning nuanced commonsense thresholds from natural-language prompts rather than relying on manual trial-and-error [2507.04293].

A plausible implication is that ARL shifts relational specification from a static engineering artifact to a runtime-adaptive interface between natural-language scene descriptions and executable geometric constraints. In AutoLayout, this interface is directly tied to the system’s broader objective of generating layouts that are both physically plausible and semantically faithful.

## 2. Internal organization and relation representation

ARL is structured into three parallel sub-libraries: Anchoring Relations, Relative Relations, and Alignment Relations. Each entry contains five elements: a symbolic name and type, a natural-language definition, a Relative Position Coordinate (RPC) template, a constraint function $f_r(\mathbf{C}) \to [0,1]$, and a Boolean validation function $\mathrm{is\_}r(\mathbf{C}) \to \{\mathtt{true}, \mathtt{false}\}$ [2507.04293].

The three sub-libraries organize relations by arity and functional role. Anchoring relations express placement of an object with respect to a boundary or support region, exemplified by $\mathit{anchor}(o_i,\mathcal{B})$. Relative relations encode binary dependencies such as $\mathit{left\_of}(o_i,o_j)$ and $\mathit{on\_top\_of}(o_i,o_j)$. Alignment relations capture multi-object regularities such as $\mathit{align\_x\_axis}(o_i,o_j,o_k)$ [2507.04293].

Internally, each relation is represented as both a JSON-like record and a small Python module. At runtime, these modules are loaded and their functions are invoked by the fast optimizer for fitness evaluation and by the self-validation checker for consistency testing. This dual representation is important: the JSON-like record provides a structured symbolic description, while the Python module provides executable semantics [2507.04293].

The following table summarizes the sub-library structure.

| Sub-library | Example relation | Role |
|---|---|---|
| Anchoring Relations | $\mathit{anchor}(o_i,\mathcal{B})$ | Placement relative to a boundary or support region |
| Relative Relations | $\mathit{left\_of}(o_i,o_j)$, $\mathit{on\_top\_of}(o_i,o_j)$ | Pairwise topological constraints |
| Alignment Relations | $\mathit{align\_x\_axis}(o_i,o_j,o_k)$ | Multi-object alignment constraints |

This representation makes ARL simultaneously symbolic and operational. The symbolic portion supports interpretability and prompting, while the executable portion allows direct participation in optimization and closed-loop correction.

## 3. LLM-based creation, prompting, and self-repair

ARL is constructed and updated through three classes of prompts, each wrapped with special tags so that they can be parsed and re-sent to GPT-4o or another large language model. The first is the relation definition prompt, whose purpose is to transform an incomplete relation name into a well-formed dictionary entry containing type, definition, and RPC. The paper’s example asks the model to complete a relation named `right_above_of` using reference entries and to output strictly formatted JSON [2507.04293].

The second is the constraint function generation prompt. Given a relation name, its natural-language definition, and a scene snippet, the model is asked to write a Python function such as `def on_top_of(objs, …):` that returns a float in $[0,1]$ measuring how well bounding boxes satisfy the relation. The third is the validation function generation prompt, which similarly requests a Boolean function such as `def is_on_top_of(objs): …` that returns `True` or `False`. These functions are enclosed within tags so that the fast system can load them on demand [2507.04293].

A central feature of ARL is runtime self-repair. Whenever, during layout optimization, a constraint or validation fails by raising an exception or by always returning $0$ or `false`, the fast system traps the error and re-issues the relevant prompt to repair the Python snippet. The updated code is then reloaded into the corresponding ARL entry [2507.04293].

This mechanism makes ARL adaptive in a stronger sense than merely adjusting thresholds. It can also regenerate executable implementations of relations during optimization. The paper explicitly contrasts this with hard-coded rules, which would require manual re-tuning for every new scene style [2507.04293].

## 4. Mathematical formalization

ARL is formalized over a set of objects $\mathcal{O}=\{o_1,\dots,o_n\}$, a set of 3D pose parameters $\mathbf{C}=\{c_1,\dots,c_n\}\subset\mathbb{R}^6$, and a set of chosen topological relations $\mathcal{R}=\{r_1,\dots,r_m\}$. A relation $r(o_i,o_j)$ is represented either as a discrete function
$$
r:\mathbb{R}^6 \times \mathbb{R}^6 \to \{0,1\}
$$
or, in soft form, as
$$
f_r:\mathbb{R}^{6n}\to [0,1].
$$
These definitions place ARL at the interface between symbolic relation selection and continuous geometric evaluation [2507.04293].

During fine-grained grounding, the overall fitness of a candidate layout $\mathbf{C}$ is defined as
$$
\mathrm{Fitness}(\mathbf{C}) \;=\; \lambda_{\text{phys}} F_{\text{phys}}(\mathbf{C}) \;+\; \lambda_{\text{sem}} \sum_{r\in\mathcal{R}} f_r(\mathbf{C}),
$$
where $F_{\text{phys}}$ encodes collision-free ($\mathrm{CF}$) and in-boundary ($\mathrm{IB}$) metrics. The typical weights are $\lambda_{\text{phys}}=0.5$ and $\lambda_{\text{sem}}=0.5$ [2507.04293].

ARL also includes an explicit threshold-adaptation mechanism driven by closed-loop self-validation. If
$$
\mathrm{is\_}r(\mathbf{C})=\mathtt{false}
$$
but
$$
f_r(\mathbf{C}) \approx 1,
$$
the large language model is re-prompted to increase or decrease the threshold $\tau_r$ by a small factor. The update is written as
$$
\tau_r \leftarrow \tau_r + \eta\bigl(\mathtt{target}_r - f_r(\mathbf{C})\bigr),
$$
with conservative learning rate $\eta \ll 1$ [2507.04293].

This suggests that ARL is designed not only to score layouts but also to detect mismatches between soft and hard relation semantics. The resulting closed loop links optimization, symbolic validation, and threshold revision.

## 5. Integration into the AutoLayout pipeline

ARL is integrated into both initialization and the two-stage layout synthesis process of AutoLayout. During initialization, default ARL entries are loaded from cache. For each user-specified new relation name, if the relation is not already present, the system prompts the large language model to create the dictionary entry, generate the constraint function, and generate the validation function [2507.04293].

In Stage 1, the coarse-grained generation phase, the slow system uses the Reasoning-Reflection-Generation (RRG) pipeline to produce a scene description. The fast system then generates discrete coordinates, extracts relations from the scene description using ARL, and validates those relations. If some object is never placed, the fast system updates placement iteratively until the missing set is empty [2507.04293].

In Stage 2, the fine-grained grounding phase, a population is initialized from the discrete coordinates. For each candidate in the population, the score begins with $0.5 \cdot F_{\text{phys}}(p)$, and then for each selected relation $r$, the system attempts to add $0.5 \cdot \mathrm{ARL}.f_r(p)$. If an error occurs, the system prompts the large language model to repair the constraint for $r$, reloads the relation, and resumes scoring. The population is then updated by selection, crossover, and mutation. The process terminates early if all constraints pass self-validation on the best candidate [2507.04293].

The red arrows in Figure 2 are identified as the points where any `is_r` failure triggers a threshold or code adjustment, thereby forming the closed loop. ARL therefore participates in both relation extraction and relation enforcement, rather than acting solely as a post hoc validation layer.

## 6. Relation exemplars and representational behavior

Three relation entries are described concretely. The binary relation `left_of` is defined as “Object A’s max X coordinate is at least $\delta$ units less than object B’s min X.” Its RPC is $[-1,0,0]$. The associated constraint code computes a distance-based ratio bounded by `delta_min=5` and `delta_max=50`, multiplies it by a Y-axis overlap term, and returns $0$ if the ordering condition fails. Its validation function returns whether $(B.min_x - A.max_x) > delta_min$ [2507.04293].

The binary relation `on_top_of` is defined as “A’s bottom Z is within a small $\epsilon$ above B’s top Z, and XY projections overlap by $\ge 70\%$.” Its RPC is $[0,0,1]$. The provided snippet returns the XY projection overlap when the vertical offset is within `eps=2.0`, and $0$ otherwise [2507.04293].

The ternary relation `align_x_axis` is defined as “Three objects share the same Y and Z centroids within tolerance $\tau$.” RPC is not used. The corresponding code checks whether the maximum-minus-minimum spread for these centroid coordinates remains below `tol=3` [2507.04293].

The following table summarizes these exemplars.

| Relation | Type | Key specification |
|---|---|---|
| `left_of` | Binary | RPC $[-1,0,0]$; max X of A is at least $\delta$ less than min X of B |
| `on_top_of` | Binary | RPC $[0,0,1]$; bottom Z of A near top Z of B with XY overlap $\ge 70\%$ |
| `align_x_axis` | Ternary | Shared Y and Z centroids within tolerance $\tau$ |

Because these relations are generated by a large language model, ARL can produce relations such as `right_above_of` or `near_of` with scene-appropriate thresholds simply by prompting. The paper explicitly presents this as a contrast to hard-coded rules, which would require manual re-tuning for each new scene style [2507.04293].

## 7. Empirical evidence and interpretive significance

The empirical evidence reported for ARL comes from an ablation in Table 7 of the main paper. Removing ARL and replacing it with a fixed, manually written rule set causes Collision-Free (CF) to drop from $99.7\%$ to $93.2\%$, In-Boundary (IB) to drop from $99.4\%$ to $97.5\%$, PSF (overall) to drop from $92.8\%$ to $90.9\%$, and the average number of sampling rounds to increase from $1.04$ to $1.38$ [2507.04293].

These results are interpreted in the source along three axes. First, ARL’s adaptive thresholds substantially reduce object overlap. Second, on-the-fly code repair avoids costly manual debugging. Third, the closed-loop self-validation converges in fewer iterations when ARL entries accurately capture scene-specific spatial norms [2507.04293].

At the level of the full AutoLayout system, the paper reports validation across 8 distinct scenarios and a significant $10.1\%$ improvement over SOTA methods in terms of physical plausibility, semantic consistency, and functional completeness [2507.04293]. Since ARL is specifically introduced to mitigate the limitations of handcrafted rules, a plausible implication is that part of this system-level gain depends on ARL’s ability to align natural-language spatial intent with executable geometric checks.

A common misconception would be to view ARL as merely a rule database. The formulation in AutoLayout is broader: ARL is a runtime-adaptive library whose entries include symbolic definitions, numerical scoring functions, Boolean validators, and mechanisms for large-language-model-mediated regeneration and threshold adjustment. In that sense, ARL is the relational substrate through which AutoLayout operationalizes closed-loop layout synthesis under slow-fast collaborative reasoning [2507.04293].

Source: https://www.emergentmind.com/topics/adaptive-relation-library-arl