---
title: Agent-Level Behavior Models
url: https://www.emergentmind.com/topics/agent-level-behavior-models
type: topic
---

# Agent-Level Behavior Models

Searching arXiv for recent and foundational papers on agent-level behavior models and closely related formulations.
Using arXiv search to verify the cited papers and anchor the synthesis in the current literature.
Agent-level behavior models are models in which the primary explanatory or predictive object is the behavior of an individual agent—its policy, trajectory, internal drivers, or deliberative process—rather than only aggregate system outputs. In recent work, the term denotes a surrogate policy learned from observed states and actions and exposed through a compact, locally interpretable behavior representation [2309.10346]; a framework that targets the conditional next-state distribution of each agent directly in agent-based models [2505.21426]; a behavior model executed for each individual traffic participant in closed-loop simulation [2512.05812]; and a framework that treats artificial agents as latent generative hypotheses about cognitive mechanisms [2604.27894]. Across these uses, the common theme is that behavior is modeled at the level at which an agent observes, decides, acts, and adapts.

## 1. Formal scope and core objects

A standard formulation places the agent in sequential decision-making. At time \(t\), the agent observes a state \(s_t\), selects an action \(a_t\), and transitions to \(s_{t+1}\), with actions sampled from \(\pi(a_t \mid s_t)\) and trajectories written as \(\tau = s_0, a_0, s_1, a_1, \dots, s_t, a_t\). In this setting, an agent-level behavior model is not limited to a single input-output explanation; it can instead model the policy around a state or the decision logic that plausibly produced the action [2309.10346].

For LLM-based agents, a trajectory may be represented as an ordered interaction history \(H_t = (o_0, a_0, o_1, a_1, \ldots, o_t)\), or as a sequence of temporally ordered components \(C = (C_1, C_2, \ldots, C_{2T})\), where each component is typed as \(\{\text{USER}, \text{THOUGHT}, \text{TOOL}, \text{OBS}, \text{MEMORY}\}\). This representation makes explicit that agent behavior depends not only on current context, but on the order of tool calls, memory updates, and intermediate thoughts [2601.15075].

Controlled behavioral evaluation introduces a different but compatible formal object. ABxLab defines an environment \(\mathcal{E} = \langle \mathcal{S}, \mathcal{A}, \mathcal{O}, \mathcal{T}, \mathcal{I} \rangle\), with intervention functions \(I : \mathcal{O}\to\mathcal{O}\) that alter what the agent sees before it is observed. The operational move is to intervene on the observation stream itself, so that the model of behavior is defined through systematic changes in option attributes and persuasive cues rather than only through task completion [2509.25609].

In cognitive behavioral science, the same topic is formalized as a joint probability model linking task variables, latent agent variables, observed actions, and rewards. For the Gabor task, the factorization is
\[
p(s_{1:T}, c_{1:T}, o_{1:T},d_{1:T},a_{1:T}, r_{1:T}) = \prod_{t=1}^T p(s_t)p(c_t|s_t) p(o_t|c_t)p(d_t|o_t)p(a_t|d_t)p(r_t|a_t,s_t).
\]
Here the agent is a latent construct: it generates internal beliefs and decisions, but the experimenter sees only the actions [2604.27894].

## 2. Representational families

The literature instantiates agent-level behavior models through several distinct representational choices.

| Representation | Core object | Representative use |
|---|---|---|
| Interpretable surrogate policy | Decision tree and local decision path | Behavior explanation [2311.18062] |
| Adaptive stochastic process | MCRE and higher-dimensional homogeneous Markov chain | Agent behavior prediction [1404.4960] |
| Structural decision model | Myopic utility maximization with latent motivational state | Behavioral analytics [1702.05496] |
| Learned nonstrategic choice | Weighted level-0 distribution over actions | Behavioral game theory [1609.08923] |
| Psychological or protocol-level model | “Inner parliament” or finite-state behavior specification | Human simulation and LLM agents [2511.02606], [2310.08535] |

In explanation-oriented work, the representation is a decision-tree surrogate \(\hat{\pi}\) distilled from sampled rollouts, together with a state-specific decision path \(dp = \mathrm{Path}(\hat{\pi}, s_t)\). This path is the behavior representation: it is local, compact, and interpretable because each tree node corresponds to a human-readable condition [2311.18062].

In platform settings with adaptive human or strategic agents, the behavior data are modeled as non-i.i.d. sequences generated by a Markov Chain in Random Environments. The state variables are joint behavior \(b_t\), feedback \(h_t\), and random user factors \(u_t\), with transition \(P(b_{t+1}\mid b_t,h_t)\). The key technical step is to augment the state space to \(\mathbb{M}=\{(h_t,b_t,b_{t+1}) : t\ge 0\}\), yielding a finite, time-homogeneous Markov chain that admits stationarity and uniform convergence analysis [1404.4960].

In intervention and incentive design, the model is often structural and agent-specific. The behavioral analytics framework assumes a discrete-time system with state \(x_t\), motivational state \(\theta_t\), action \(u_t\), and incentive \(\pi_t\), with dynamics \(x_{t+1}=h(x_t,u_t)\) and \(\theta_{t+1}=g(x_t,u_t,\theta_t,\pi_t)\). The myopia assumption gives
\[
u_t \in \argmax \big\{ f(x_{t+1},u,\theta_t,\pi_t)\ \big|\ x_{t+1}=h(x_t,u),\ u\in\mathcal{U}\big\},
\]
so the same individualized model supports estimation, prediction, and optimization [1702.05496].

A classical antecedent in contract theory defines agent behavior through expected payoff
\[
E_A(e)=\sum_{i=1}^n p_i(e)\,u(w_i)-v(e),
\]
with derivative \(M_A(e)=\frac{dE_A}{de}\) called the motivation function and second derivative \(Prst_A(e)=\frac{d^2E_A}{de^2}\) called the persistence function. This broadens the standard assumption that effort is always a pure loss and makes risk contract-dependent rather than a fixed trait of the agent [1107.2881].

Behavioral game theory contributes another representational lineage. Level-0 behavior is modeled as a learnable distribution over actions based on features such as maxmax payoff, maxmin payoff, minmin unfairness, and max symmetric, rather than as a uniform distribution \(\pi^{\varnothing}_{i,0}(a_i)=|A_i|^{-1}\). The recommended model is a linear weighting of features that requires the estimation of four weights [1609.08923].

Psychological and protocol-level models replace utility or state-space dynamics with explicit internal roles or explicit behavioral grammars. The psychological simulation system models a person as an “inner parliament” of agents such as Threat-Avoidance, Math-Anxiety, Goal-Pursuit, Procedural-Fluency, and Self-Efficacy [2511.02606]. The formal specification framework models an LLM agent as a finite-state machine \(\langle \mathcal{D}_S, \delta, s_0, s_{end} \rangle\) together with a declarative behavior formula expressed with operators such as `next`, `until`, and `or` [2310.08535].

## 3. Learning and inference from observations, logs, and video

A major line of work learns behavior models from observed trajectories only. In the behavior-explanation setting, the policy is distilled into a decision tree using imitation learning, specifically DAgger:
\[
\hat{\pi} = \arg\min_{\pi \in \Pi} \mathbb{E}_{s^*,a^* \sim \pi^*}[\mathcal{L}(s^*, a^*, \pi)].
\]
The explanation system is therefore independent of the agent’s internal representation and can be used for any policy from which trajectories can be sampled [2311.18062].

Agent-based-model surrogates adopt a micro-level objective rather than a system-level one. The Graph Diffusion Network learns the conditional next-state distribution of each agent directly,
\[
\mathbf{Z}^{(i)}_{t+1} \sim P_{\Theta}(\mathbf{Z}^{(i)}_{t+1}\mid \mathbf{Z}^{(i)}_t,\{\mathbf{Z}^{(j)}_t\}_{j\in N_t^{(i)}}),
\]
using a message-passing graph neural network to encode local interactions and a conditional diffusion model to capture behavioral stochasticity. Training is performed on a “ramification” dataset built from a main branch and many sibling branches, with denoising loss
\[
L(\phi,\omega) = \mathbb{E}_{i,t,\tau,\epsilon} \left[ \|\epsilon-\epsilon_\phi(\tilde{\mathbf{Z}_{t+1}^{(i)}(\tau)},\mathbf{c}_t^{(i)})\|^2 \right].
\]
The reported experiments show that the surrogate replicates individual-level patterns and forecasts emergent dynamics in Schelling’s segregation model and a Predator-Prey ecosystem [2505.21426].

Video-based behavior learning replaces simulator rollouts with persistent reconstruction. Agent-to-Sim first builds a persistent spacetime 4D representation from monocular RGBD smartphone videos, with a canonical structure \({\bf T} = \{\sigma, c, \boldsymbol{\psi}\}\) and time-varying structure \(\mathcal{D} = \{\boldsymbol{\xi}, {\bf G}, {\bf W}\}\). It then trains a hierarchical diffusion model over goal, path, and full body motion, conditioned on egocentric scene, observer, and past-motion codes. On the cat split, the main behavior metrics with \(K=16\) samples are Goal \(0.448 \pm 0.146\) m, Path \(0.234 \pm 0.054\) m, Orientation \(0.550 \pm 0.112\) rad, and Joint angles \(0.237 \pm 0.006\) rad [2410.16259].

Policy ABMs use RL as a behavior model when agents are treated as adaptive, reward-seeking decision-makers. The interface is gym-style,
\[
s_{t+1}, r_t, done = env.step(a_t), \qquad s_0 = env.reset(),
\]
and the paper studies REINFORCE, REINFORCE with baseline, Actor-Critic, and a multi-agent actor-critic adaptation in the minority game and an influenza-transmission ABM. In the flu case, after about 750 training iterations with actor-critic, the RL agent improves its outcomes and on average outperforms the default behavioral model by a few percentage points over 100 seasons [2006.05048].

A different simulation-based inference scheme is agent-centric Monte Carlo cognition. Here the agent’s “thinking” is another ABM: at each tick, the agent initializes a cognitive model from local observations, runs a settable number of rollouts for each candidate action, computes discounted reward based on change in energy, and executes the action with the highest mean reward. In the Wolf-Sheep Predation example, increasing the number of rollouts improves performance monotonically, and rollout length is most effective up to about three ticks [1807.10847].

## 4. Explanation, attribution, and procedural analysis

One prominent use of agent-level behavior models is natural-language explanation grounded in a surrogate policy. The explanation pipeline is:
\[
\text{trajectory data} \rightarrow \text{decision tree surrogate} \rightarrow \text{local decision path} \rightarrow \text{LLM explanation}.
\]
The prompt has four parts: a description of the environment and action semantics, a description of what the behavior representation means, in-context examples, and the specific behavior representation and action to explain. In an IRB-approved study with 32 participants, explanations generated from the behavior representation were significantly preferred over template-based explanations (\(p < 0.05\), one-tailed binomial test), and participants did not prefer human explanations over the proposed method. In a hallucination analysis over 30 generated explanations, the behavior-representation method produced significantly fewer hallucinations than both baselines [2309.10346].

General agentic attribution shifts the question from failure localization to action explanation regardless of task outcome. At component level, it computes
\[
V_i = \log p_{T_\theta}(a_T \mid C_{\le i}), \qquad
g_i = V_i - V_{i-1},
\]
so a large positive temporal gain marks a decision driver. At sentence level, it scores sentences with
\[
D_{i,j} = \text{Drop}(s_{i,j}) + \text{Hold}(s_{i,j}),
\]
where probability drop measures necessity and probability hold measures sufficiency. Using human-annotated ground truth sentences, the default Prob. Drop&Hold method achieves Hit@1 \(= 0.944\), Hit@3 \(= 1.000\), and Hit@5 \(= 1.000\), outperforming LOO, ContextCite, and Saliency within the same hierarchical framework [2601.15075].

Coding-agent analysis treats an agent run as a trajectory through a code state-space. The Simple Strands Agent paper formalizes the “intent-execution gap” between what the model intends and what the harness executes, defines repository states \(x:\mathcal{P}\to \Sigma^\ast \cup \{\bot\}\), and measures live progress with a recall-style divergence
\[
D_i(t)=\min_{x^\star\in \mathcal{S}_i} \left( 1-\frac{|\phi_i(x_i(t))\cap \phi_i(x^\star)|}{|\phi_i(x^\star)|} \right).
\]
Across SWE-Bench-Verified, SWE-Bench-Pro, and Terminal-Bench-2, the study analyzes about 138k trajectories. Its core empirical claim is that similar pass@1 scores can conceal very different edit frequency, testing activity, phase composition, and backtracking behavior [2606.17454].

A complementary line represents trajectories themselves as procedures. “Trajectories as programs” induces an emergent vocabulary over action sequences with BPE-style merging, selects vocabulary size by V-measure, and reports a stable choice at \(K = 192\) with V-measure \(= 0.644\). A probe over procedural fingerprints attributes unseen trajectories to the correct agent at 85.7% accuracy, against an 11.1% chance baseline for ten agents. ProcGrep, the structural querying library introduced in the same work, attains mean F1 \(= 1.000\) with latency \(1.1\,\mu s\) on its benchmarked structural queries, whereas Claude Sonnet 4.6 obtains F1 \(= 0.278\) with latency \(1.71\) s and GPT-4o obtains F1 \(= 0.230\) with latency \(0.66\) s [2606.16988].

This suggests that explanation is no longer restricted to post hoc verbal rationalization. In this literature it includes local surrogate policies, historical attribution over tool and memory traces, and procedural fingerprints over full trajectories.

## 5. Simulation, intervention, and human-centered applications

In multi-agent decision support, agent-level behavior models can drive personalized interventions. The behavioral analytics framework organizes the workflow into three steps: develop a behavioral model, estimate behavioral model parameters and predict future decisions, and optimize a set of costly incentives. The single-agent 2SSA algorithm first computes a MAP estimate of \((x_0,\theta_0)\) and then optimizes future incentives; the multi-agent ABMA algorithm decomposes the global problem into one MAP estimation per agent, one MILP per agent per discrete incentive type, and one master ILP. In a clinically supervised mobile weight-loss intervention, the framework maintains efficacy comparable to the original trial while reducing the required clinical-visit resources by up to about 60% [1702.05496].

Driving simulation treats the behavior model as the policy executed for each traffic participant at every simulation step. The instance-centric design represents each traffic participant and each map element in its own local coordinate frame, uses a query-centric symmetric context encoder with relative positional encodings, and learns behavior with Adversarial Inverse Reinforcement Learning plus an adaptive reward offset \(c(n)=\tilde{r}_\mathrm{target} - \tilde{r}_\mathrm{mean}(n)\). The paper reports that the small model with only 59K parameters still outperforms all baselines on most metrics, that training time is reduced by up to 66.4% compared to agent-centric baselines, and that peak inference throughput reaches about 88k and 358k ISPS, up to 13.2× improvement over agent-centric baselines [2512.05812].

Human behavior simulation uses a more explicit internal architecture. The multi-agent psychological simulation system models a person as an “inner parliament” of psychological agents corresponding to constructs such as Threat-Avoidance, Math-Anxiety, Spatial-Reasoning, Goal-Pursuit, Procedural-Fluency, and Self-Efficacy. Behavior generation proceeds through initial positions, debate and adjustment over several rounds, and final synthesis. The system’s “Peek Into the Brain” view exposes the visible transcript of internal deliberation, and the paper frames the goal not as task optimality but as psychological authenticity in teacher training, psychological research, and related simulations [2511.02606].

Policy-focused ABMs offer another application domain. In the minority game and influenza-transmission case studies, RL behavioral models are used as adaptive, reward-seeking or reward-maximizing agents. The experiments show that such agents can learn to outperform the default adaptive behavioral models in the two ABMs examined, while synchronization analysis in the flu setting finds no strong synchronization, with correlations capped at about 10% magnitude [2006.05048].

## 6. Evaluation, governance, and recurrent tensions

Behavioral evaluation frameworks argue that competence alone is insufficient. ABxLab turns websites into behavioral experiments by manipulating prices, ratings, order, and synthetic nudges in a 2AFC shopping environment. The benchmark runs over 80,000 experiments across 17 models, over about 2.5B tokens and about 400k requests. In the Original condition, higher ratings increased selection probability for 14 of 17 models with effects typically 30–80 pp, and the strongest effect was 81.2 pp for o4-mini. Price effects in the Original condition were about 15–34 pp for 13 of 17 models, while nudges shifted choices by about 10–60 pp on average when price and ratings were matched. Humans were far less sensitive overall, with order about 4 pp, rating about 5 pp, price about 9.4 pp, and nudge about 9.9 pp [2509.25609].

Governance-oriented work extends the unit of analysis from local choice to the full lifecycle of network action. The “Network Behavior Lifecycle” decomposes behavior into Target Confirmation, Information Gathering, Reasoning Process, Decision Mechanism, Action Execution, and Feedback Acquisition. The proposed A4A paradigm uses agents to regulate, monitor, or evaluate other agents, while HABD compares humans and agents along five dimensions: decision mechanism, execution efficiency, intention-behavior consistency, behavioral inertia, and irrational patterns. In the red-team case, PentAGI required 2,000,000 GPT-4o tokens in zero-shot mode and 500,000 tokens with structured CoT prompts; in the blue-team case, the defensive coding comparison reports total time 68.2 s for the agent and 615 s for the human [2508.14415].

Behavior can also be engineered declaratively rather than inferred post hoc. The formal specification framework lets a user define states and prompts, write a behavior formula such as
```lisp
(next
  Ques
  (until
    (next Tht Act Act-Inp Obs)
    Final-Tht)
  Ans)
```
for ReACT-like behavior, and compile the specification into a decoding monitor that validates, truncates, and repairs generated text through valid state prefixing. The same framework reproduces ReACT, ReWOO, Reflexion, direct answering, and chain-of-thought, and introduces PASS, a Plan-Act-Summarize-Solve agent. On HotpotQA, TriviaQA, and GSM8K, PASS is best among agent-based methods on HotpotQA and TriviaQA, while ReACT is best among agent-based methods on GSM8K [2310.08535].

Several papers also make the limits of the field explicit. General agentic attribution argues that prior work predominantly focuses on failure attribution and is therefore blind to a correct but poorly justified answer, memory-induced bias, tool-conditioned hallucination, or overgeneralization from a prior success case [2601.15075]. The coding-agent trajectory work argues that pass@1 is necessary but insufficient because similar solve rates can conceal large differences in trajectory quality and because public containers can leak future git history [2606.17454]. The psychological simulation paper is conceptual and architectural and does not provide a detailed mathematical formalism, explicit equations, or pseudo-code for the deliberation algorithm, while the governance paper describes HABD as qualitative and preliminary rather than fully formalized [2511.02606], [2508.14415].

Taken together, these formulations indicate that agent-level behavior models are not a single model class. They are a research program centered on the behavior of the individual agent as the unit of explanation, prediction, intervention, or governance. This suggests a broad convergence: surrogate policies, stochastic process models, utility-based decision systems, procedural trajectory models, and psychologically grounded simulators are all being used to ask the same technical question—how an agent behaves, why it behaves that way, and how that behavior should be evaluated or controlled.

Source: https://www.emergentmind.com/topics/agent-level-behavior-models