---
title: Behavior Consistency Reward (BehR)
url: https://www.emergentmind.com/topics/behavior-consistency-reward-behr
type: topic
---

# Behavior Consistency Reward (BehR)

Behavior Consistency Reward (BehR) is a step-level reward for training and evaluating text-based world models by measuring whether a predicted state preserves the action tendency of a downstream agent rather than merely matching the surface form of the real state. In the formulation introduced in "Beyond State Consistency: Behavior Consistency in Text-Based World Models" [2604.13824], BehR operationalizes *functional consistency* through the likelihood of a logged next action under a frozen Reference Agent, and is optimized through a GRPO-based post-training loop for world models. The same source positions BehR as an alternative to single-step metrics such as Exact Match, while related work frames behavior consistency more broadly as a reward-smoothing or decision-preservation principle in RLHF and multi-agent RL [2506.09096] [2312.05783].

## 1. Functional consistency and the definition of BehR

The ideal target underlying BehR is *functional consistency*. Let $\pi$ be a fixed downstream agent, and let $h_t = (s_0, a_1, s_1, \ldots, a_t)$ denote the history. A world model $\hat H$ is functionally consistent with the real environment if, for every step $t$,
$$
\pi(a \mid h_t, s_{t+1}) = \pi(a \mid h_t, \hat s_{t+1}) \quad \text{for all actions } a.
$$
This condition states that replacing the real next state $s_{t+1}$ with the predicted next state $\hat s_{t+1}$ should leave the downstream agent’s action distribution unchanged [2604.13824].

In practice, the full distribution $\pi(\cdot \mid \ldots)$ is often inaccessible, especially when the downstream agent is a black box. BehR replaces this intractable objective with a tractable probe based on the logged next action $a^*_{t+1}$. A frozen Reference Agent $\pi_{\mathrm{ref}}$ is used to score that action under both the real state and the predicted state. Given context $c = (h_t, \text{state})$ and logged next action $a^*_{t+1}$, the reference model returns the mean per-token log-probability
$$
\bar\ell(c, a^*) \;\triangleq\; \frac{1}{|a^*|} \cdot \log \pi_{\mathrm{ref}}(a^* \mid c).
$$
The real and predicted scores are then
$$
\bar\ell_{\mathrm{real}} = \bar\ell((h_t, s_{t+1}), a^*_{t+1}), \qquad
\bar\ell_{\mathrm{pred}} = \bar\ell((h_t, \hat s_{t+1}), a^*_{t+1}),
$$
and BehR is defined as
$$
R_{\mathrm{beh}} = - \left| \bar\ell_{\mathrm{pred}} - \bar\ell_{\mathrm{real}} \right|.
$$
This reward lies in $(-\infty, 0]$ and is maximized at zero when the predicted state preserves the reference-agent likelihood of the logged action exactly. The paper also notes that a bounded transform such as $\exp(-|\Delta|)$ can be used in practice so that $R_{\mathrm{beh}} \in (0,1]$, while preserving the same core objective [2604.13824].

The central distinction is therefore between *state similarity* and *behavioral faithfulness*. BehR asks whether the generated state leads the agent to make the same decision, not whether the text simply looks similar to the real state. This suggests a shift from reconstruction-oriented training to decision-preserving training.

## 2. Derivation and intuition

The derivation of BehR begins from the ideal requirement that the downstream policy behave identically under the real and predicted next states. Because this full equality is generally unavailable when $\pi$ is a black box, the logged next action is used as a probe of the decision boundary. If $\hat s_{t+1}$ preserves the likelihood of $a^*_{t+1}$ relative to $s_{t+1}$, then the world model has preserved at least one decision-critical aspect of the environment state [2604.13824].

The resulting quantity is dense and step-level. The paper contrasts this with binary task success and with token-level reconstruction metrics. Under this view, the absolute gap
$$
\left| \bar\ell_{\mathrm{pred}} - \bar\ell_{\mathrm{real}} \right|
$$
measures how much the candidate state perturbs the downstream action preference of the frozen judge. Smaller gaps indicate greater behavioral faithfulness. A plausible implication is that BehR privileges latent state properties that matter causally for action selection, even when they are not the dominant contributors to surface-form similarity.

This design also clarifies what BehR is not. It is not a direct estimate of environment correctness in a semantic or factual sense, and it is not a full imitation of the downstream policy. Rather, it is a proxy for functional consistency defined through a single logged action under a surrogate agent. The paper explicitly motivates this substitution by the inaccessibility of the full action distribution and by the need for a tractable reward [2604.13824].

A related but distinct formulation appears in "Intra-Trajectory Consistency for Reward Modeling" [2506.09096], where consistency is imposed over adjacent prefixes in autoregressive text generation. There, the Bayesian motivation is that as the generator probability of a continuation approaches one, the variance of the reward of a prefix around the reward of its extension vanishes; in words, “if the model is almost certain that extension $y_{m+1:n}$ follows $y_{1:m}$, then their rewards should be almost the same.” The paper enforces this locally between adjacent prefixes rather than over arbitrary continuations. This provides a closely related interpretation of BehR as a local smoothness principle specialized to sequence generation [2506.09096].

## 3. Optimization with GRPO

BehR is integrated into an RL-style world-model training loop through Group Relative Policy Optimization (GRPO). The world model $\hat H_\theta$ is treated as a policy $\pi_\theta$ over next states conditioned on $(h_t, a_t)$, and the training objective is to maximize expected BehR across an offline dataset of tuples $(h_t, s_{t+1}, a^*_{t+1})$ [2604.13824].

The procedure described in the paper is as follows. For each minibatch of prompts $h_t$, the model samples $n$ candidate next states $\{\hat s_{t+1}^i\}_{i=1\ldots n}$ from $\hat H_\theta$ using temperature $T>1$ to increase diversity, with the main experiments using $n=5$ and $T=1.3$. For each prompt, the real reference score $\bar\ell_{\mathrm{real}}$ is cached to avoid duplicate calls, and each candidate state receives a BehR score
$$
R^i = -\left| \bar\ell((h_t,\hat s^i),a^*) - \bar\ell((h_t,s),a^*) \right|.
$$
Within each prompt group, the $n$ rewards are normalized by subtracting the group mean and dividing by the group standard deviation. The GRPO objective is then
$$
L(\theta) = - \mathbb{E}_{h_t,\hat s^i}\bigl[w^i \cdot \log \pi_\theta(\hat s^i \mid h_t, a_t)\bigr],
$$
where $w^i$ is the normalized BehR, together with an auxiliary KL penalty $\beta \cdot \mathrm{KL}(\pi_\theta \| \pi_{\mathrm{SFT}})$ to prevent drift. The main experiments use $\beta = 0.001$ and run under the VeRL/GRPO framework with batch size $32$, micro-batch per GPU, and FSDP plus tensor parallelism [2604.13824].

The paper reports that the Reference Agent $\pi_{\mathrm{ref}}$ is frozen throughout training, with Qwen3-8B given as an example. This avoids a moving target. It also states that the exponential reward mapping $\exp(-|\Delta|)$ is used in the main experiments, while linear and Cauchy variants are also supported [2604.13824].

This optimization scheme positions BehR as a post-training signal layered on top of an SFT-initialized world model rather than as a replacement for supervised pretraining. A plausible implication is that BehR is intended to redirect generation toward decision-critical fidelity while preserving the broad linguistic and structural competence supplied by SFT.

## 4. Empirical findings in WebShop and TextWorld

The principal empirical study evaluates BehR on WebShop and TextWorld. WebShop is described as an e-commerce simulator where states are open-ended product catalog pages and actions are `search[keywords]` or `click[ASIN]`; success is purchasing the correct item within 50 steps. TextWorld is a structured text-adventure generator with rooms, objects, and subgoals; actions are natural-language commands such as “open chest” or “go north,” and success is completing all subgoals within 50 turns [2604.13824].

The base world models are Word2World SFT checkpoints on Qwen2.5-7B and LLaMA3.1-8B. BehR-WM denotes the SFT checkpoint followed by BehR-GRPO post-training. RL ablations include F1-WM, which uses GRPO with token-level F1 reward, FactR, which targets structured factual correctness, and BehR-WM. Evaluation agents include Qwen3-8B, Qwen3-32B, GPT-4o, and GPT-5, all at zero temperature [2604.13824].

On single-step prediction, BehR post-training preserves Exact Match on WebShop, with $79.05 \rightarrow 79.19\%$, and improves on TextWorld, with $64.7 \rightarrow 73.1\%$, in three out of four model-backbone setups. On task-level functional consistency, measured through the pairwise consistency ratio $\mathrm{CR}_{pw}$, BehR yields larger gains. For each held-out initial task, the agent acts in the real environment, then in the world model, and then the world-model-generated action sequence is replayed in the real environment; $\mathrm{CR}_{pw}$ is defined as the fraction of Real-successful tasks that remain successful under W2R. In WebShop, using the Qwen2.5-7B base world model and a Qwen3-8B agent, $\mathrm{CR}_{pw}$ rises from $0.345$ for SFT to $0.483$ for BehR. The paper reports similar gains for Qwen3-32B and GPT-4o. In TextWorld, the gains are smaller but consistent in non-ceiling regimes, with one reported setting improving from $0.678$ to $0.730$ [2604.13824].

The same study also examines calibrated surrogate evaluation. In offline “WM as evaluator” mode, W2W-WM is reported to greatly overestimate weak agents, with false positives up to $42.5\%$ on TextWorld Qwen3-0.6B. BehR-WM reduces false-positive rate by $3\times$–$5\times$, raising task-level agreement between world-model and real-environment evaluation from approximately $57\%$ to $90\%$ in the weakest cases [2604.13824].

For inference-time lookahead planning with $K=5$, agents propose $K$ candidate actions, the world model simulates next states, and the agent picks the best outcome. Lookahead with SFT-WM already improves success rate by $+7$–$9$ percentage points, and BehR-WM yields marginal extra gains, reported as $+2$ percentage points in some cases. The paper characterizes these results as preliminary and points toward safer planning [2604.13824].

These findings support a narrow but important claim: optimizing a behavior-aligned signal can improve long-term alignment even when single-step metrics move only slightly. They do not establish that BehR dominates all alternative objectives in every regime; the source explicitly notes less movement in near-ceiling regimes [2604.13824].

## 5. BehR in relation to reward consistency in RLHF

A closely related notion appears in reward modeling for RLHF. "Intra-Trajectory Consistency for Reward Modeling" introduces an intra-trajectory consistency regularization that propagates response-level supervision across adjacent processes in a response trajectory using generation probabilities [2506.09096]. The generator’s next-token probability is written as
$$
p_t \equiv p_{\theta_{(g)}}(x_t \mid x_{<t}),
$$
and the reward model’s calibrated output on the prefix $x_{1:t}$ is
$$
R_\phi(x_{1:t}) \equiv \hat r(x, x_{1:t}).
$$
Under a Bradley–Terry calibration, for a chosen response $y^w$,
$$
\hat r(x,y^w_{1:m}) = \sigma\!\left( \theta_r(x,y^w_{1:m}) - \frac{1}{|y^l|}\sum_{k=1}^{|y^l|}\theta_r(x,y^l_{1:k}) \right),
$$
with an analogous definition for a rejected prefix. The regularizer constructs cross-entropy terms over adjacent prefixes whose weights depend on both the next-token generation probability and the model’s confidence in the other prefix’s reward, and combines them with the standard Bradley–Terry loss through
$$
\mathcal{L}_{\mathrm{total}} = (1-\alpha)\mathcal{L}_{bt} + \alpha \mathcal{L}_{reg}.
$$
The paper states that this is “exactly a data-driven instantiation” of the RLHF idea of augmenting scalar outcome reward with a local smoothness or consistency term over likely adjacent transitions, and summarizes the induced behavior-consistency reward as
$$
r_{\mathrm{beh}}(x,y_{1:k-1},y_{1:k}) \propto - \left|R_\phi(x,y_{1:k}) - R_\phi(x,y_{1:k-1})\right| \times p_{\theta_{(g)}}(y_k \mid x, y_{1:k-1}).
$$
Its practical take-aways are to weight the consistency penalty by next-token probability, calibrate process rewards, focus on adjacent prefixes, and tune the balance $\alpha$ so as not to wash out the main preference signals [2506.09096].

Empirically, that regularizer improves outcome reward modeling and downstream alignment. On RewardBench with a Gemma-2B-it backbone and 40K samples from Unified-Feedback, baseline GRM average accuracy rises from $73.0\%$ to $75.8\%$ with intra-trajectory consistency; with 400K samples, $73.2\% \rightarrow 75.7\%$. Using Skywork plus Unified-Feedback and Llama3-8B-instruct with an EMA of process rewards, GRM-average improves from $86.8\%$ to $89.1\%$. In RLHF, a DPO policy evaluated by a gold scorer improves from $0.676$ to $0.678$, and the “win” rate versus GRM rises from $47.3\%$ to $50.6\%$. In inference-time Best-of-$N$ on math tasks with a Qwen-1.5B-Instruct reward model, Mistral v0.3 pass@$8$ improves from $12.6\%$ to $14.6\%$, and Llama-3-8B-Instruct pass@$8$ improves from $49.2\%$ to $51.2\%$ [2506.09096].

A separate RLHF perspective on consistency appears in "The Trickle-down Impact of Reward (In-)consistency on RLHF" [2309.16155]. That work defines *response consistency* and *instruction consistency* on a Contrast Instructions benchmark:
$$
\mathcal{C}_{\mathrm{res}} =
\frac{1}{|\mathcal D|}\sum_{(I^A,I^B,r^A,r^B)\in\mathcal D}
\Big[\mathcal R_\theta(I^A \circ r^A) > \mathcal R_\theta(I^A \circ r^B)\Big],
$$
$$
\mathcal{C}_{\mathrm{ins}} =
\frac{1}{|\mathcal D|}\sum_{(I^A,I^B,r^A,r^B)\in\mathcal D}
\Big[\mathcal R_\theta(I^A \circ r^A) > \mathcal R_\theta(I^B \circ r^A)\Big].
$$
The same paper reports that a baseline LLaMa-7B reward model trained with the standard ranking objective is near random on this benchmark, with $\mathcal{C}_{\mathrm{res}} \approx 53.6\%$ and $\mathcal{C}_{\mathrm{ins}} \approx 49.4\%$, while ConvexDA and RewardFusion improve those values modestly; humans score approximately $82\%$ on both metrics. It also reports downstream RLHF gains, including human pairwise preference above $60\%$ for the consistency-trained model and an increase in usefulness judgments from approximately $50.9\%$ to $57.8\%$ [2309.16155].

Taken together, these results suggest that BehR is part of a broader consistency-oriented family of objectives: one branch targets world-model state prediction through downstream action likelihoods, and another targets reward modeling through local reward smoothness or semantic coherence. The shared theme is preservation of decision-relevant structure rather than optimization of surface-form fidelity alone.

## 6. Broader consistency-reward formulations and open distinctions

The phrase *behavior consistency* is not unique to text-based world models. In "DCIR: Dynamic Consistency Intrinsic Reward for Multi-Agent Reinforcement Learning" [2312.05783], behavior consistency is defined as the divergence in output actions between two agents given the same observation. If agent $i$ observes $o_t^i$ and produces action distribution
$$
u_t^i = \pi_i(o_t^i),
$$
and $u_t^{ij} := \pi_j(o_t^i)$ denotes the distribution that agent $j$ would produce on that same observation, then the inconsistency metric is
$$
C_{i,j,t} = D_{KL}(u_t^{ij}\|u_t^i).
$$
The dynamic consistency intrinsic reward is
$$
r_{DCIR}^{i,t} = \sum_{j\in\mathcal N(i)} \alpha_j^i \cdot C_{i,j,t},
$$
and is combined with extrinsic reward through
$$
r_{\mathrm{proxy}}^{i,t} = r_{ex}^t + \beta \cdot r_{DCIR}^{i,t}.
$$
Here the sign of $\alpha_j^i$ determines whether the agent is rewarded for reducing divergence or increasing divergence, and a Dynamic Scale Network produces these coefficients from the joint observation [2312.05783].

This multi-agent formulation differs from BehR in target and mechanism. BehR compares *real-state* and *predicted-state* action likelihoods under a frozen Reference Agent. DCIR compares *

Source: https://www.emergentmind.com/topics/behavior-consistency-reward-behr