Behavior World Model
- Behavior World Models (BWMs) are internal models that predict and simulate agent behavior, adapting to environmental interactions and dynamically adjusting actions.
- BWMs are used in various fields, including robotics, social interactions, and human-motion prediction, to model complex behavioral dynamics and interactions over time.
- BWMs can incorporate structured representations, probabilistic dynamics, and human-like mental states, but they often lack direct physical simulations.
A Behavior World Model (BWM) is a predictive internal model that represents behaviorally relevant aspects of agents, environments, actions, observations, and temporal evolution, and uses that representation for behavior prediction, simulation, planning, control, or evaluation. The term encompasses several complementary formulations: symbolic separation of structure, events, and chronologies; probabilistic task-phase estimation coupled to behavior selection; structured representations of behavioral attributes; latent models of human motion; action-conditioned visual simulators for robot learning; physical–mental models of situated agents; and interactive models that predict how embodied behavior evolves under environmental constraints. A BWM is therefore broader than a policy, a behavior embedding, or a next-frame predictor: it should connect observations and actions to state transitions, future consequences, behavioral alternatives, and, where relevant, the beliefs, goals, norms, and intentions that generate behavior.
1. Conceptual foundations and scope
The conceptual basis of BWM research is the separation of three modeling levels: static structure, dynamic events, and behavioral chronology. Static modeling describes entities, internal structures, operations, and possible flows. Dynamic modeling identifies events and their hierarchical organization. Behavioral modeling specifies which temporal arrangements of events are acceptable. This distinction is developed through the thinging machine (TM) framework, whose basic ontological element is the thimac, simultaneously a “thing” and a “machine” (Al-Fedaghi, 2020).
In TM, a static model describes possible structures and flows rather than one realized trajectory. A dynamic model converts operations into time-indexed events; an event is defined as “a thimac that includes a time.” Behavior is defined as “the chronology of events in the TM model.” A chronology may contain alternative, repeated, concurrent, or partially ordered events rather than a single deterministic sequence. This separation prevents a flow diagram from being interpreted simultaneously as an ontology, an execution trace, and a behavioral policy.
The five generic TM operations are:
- Create: bring a thing into existence.
- Process: transform a thing.
- Release: mark a thing as ready to leave a machine.
- Transfer: transport a thing outside a machine.
- Receive: allow a thing to enter and be accepted by a machine.
These operations provide a symbolic vocabulary for objects, resources, signals, actions, observations, state changes, and interactions. A thimac may contain subthimacs, memory, and triggering relations that do not correspond directly to flow. Event hierarchies can represent primitive interactions, actions, episodes, scenes, and larger task structures.
The TM perspective is conceptual and symbolic rather than predictive in the modern machine-learning sense. It does not by itself define probabilistic dynamics, continuous physics, uncertainty models, learned representations, or inference algorithms. A BWM derived from TM must therefore augment the symbolic ontology with explicit state-transition semantics, stochasticity, continuous dynamics, agent models, and learning mechanisms.
A useful distinction among neighboring systems is:
| System | Primary modeled object |
|---|---|
| Behavior model | An agent’s action distribution |
| Policy | A mapping from observations or histories to actions |
| World model | Environment dynamics and action-conditioned consequences |
| Behavior World Model | Behaviorally relevant state, dynamics, action alternatives, and consequences |
A policy predicts or selects what an agent does. A conventional world model predicts what happens under actions. A BWM combines these concerns by modeling the state variables and temporal processes that make behavior possible, including the interaction between agent behavior and environmental response.
2. State, events, behavior, and chronology
A BWM requires a distinction between action schemas, executed actions, events, observations, and state transitions. An action schema describes a possible operation, such as transferring food or picking up an object. An action execution is a particular occurrence involving an agent, object, and time. An observed event is evidence that the execution occurred. The resulting transition changes the modeled world state.
The restaurant and food-service examples illustrate why behavior should be represented as a set of possible trajectories rather than a single script. Ordering, preparing, serving, eating, and paying can occur in different acceptable arrangements. In one service, preparation precedes ordering; in another, payment occurs at the end; in a buffet, preparation may precede ordering and payment may occur before eating. A shared static structure can therefore support multiple behavioral chronologies without duplicating the underlying entity and flow model (Al-Fedaghi, 2020).
This perspective is compatible with partial-order behavior. A constraint such as indicates that event must precede event , while unrelated events may remain concurrent or independently ordered. A BWM may represent:
- legal action sequences;
- workflows;
- temporal constraints;
- event dependencies;
- alternative policies;
- repeated events;
- hierarchical episodes;
- partial-order plans;
- constraints on what can happen next.
The dynamic and behavioral distinction is also central to robotic control. A task may have nominal phases such as free space, pre-grasp, contact, grasped, and completed, but disturbances can cause those phases to occur in an unexpected order. Baum and Brock model manipulation as online inference over discrete task states, each associated with an observation distribution, a controller, and a belief state. A Hidden Markov Model combines force–torque and visual evidence, and the controller associated with the maximum-posterior state is executed (Baum et al., 2022).
That system uses a weakly structured transition matrix with and for . The matrix expresses persistence without enforcing a nominal task sequence. Sensor evidence can move the belief directly to a contact or recovery state when a disturbance violates the expected order. In reported drawer-opening experiments, the reactive method achieved $10/10$ success under both light and strong disturbances, while the sequential baseline achieved zero successes under strong disturbance. The experiments used ten trials per condition.
This formulation is a minimal BWM: it represents latent task state and state-conditioned behavior but does not learn rich future observation dynamics or long-horizon consequences. Its principal contribution is the integration of continual state estimation and behavior selection. Interactive perception is particularly important: the robot deliberately applies forces or maintains contacts that make task states more distinguishable.
3. Representational forms of behavioral structure
A BWM may represent behavior symbolically, graphically, probabilistically, or through structured latent variables. “On the Expressive Power of Behavior Structure” proposes the Behavioral Molecular Structure (BMS), in which behavioral attributes function as atoms and relations among co-occurring attributes function as bonds (Wang et al., 2023).
A behavior is represented as a collection of measurable attributes, such as location, time, transaction type, account identity, topic, weapon, or user. A behavior-specific graph contains attribute-value nodes, typed edges, and node features. Domain-specific meta-rules determine which attributes are connected. Repeated co-occurrence strengthens relations in the global behavioral graph, while each observed behavior corresponds to a subgraph.
BMS uses relational graph convolution and average pooling to construct behavior embeddings. In crime detection, categorical values are represented with BERT embeddings, projected to 128 dimensions, processed by an RGCN, and classified with a three-layer fully connected network. In recommendation, feature-specific embedding tables produce 64-dimensional node embeddings. In fraud generation, heterogeneous transaction graphs are modeled with GraphVAE.
The paper’s expressive-power argument contrasts attribute-value representations with structure-based representations. If behavior has dimensions and each dimension has at most possible values, an ordinary representation has an upper bound of combinations. A structure-based representation with pairwise connectivity can have an upper bound of 0; for binary connectivity this becomes 1. The authors caution that this is a preliminary counting argument: a larger theoretical configuration space does not guarantee better task performance, because real graphs are constrained, sparse, noisy, and not necessarily learnable.
Empirically, BMS ranks first on five reported crime-classification indicators and second on the other two, although the supplied results do not provide their numerical values. Its advantage increases as behavioral dimensionality grows beyond approximately 9–12 attributes. However, XGBoost nearly matches BMS, demonstrating that theoretical expressive power does not automatically translate into superior prediction.
In user-behavior prediction, BMS improves some sequential recommendation models but harms others. It improves GRU4Rec, STAMP, and FEARec while reducing performance for FPMC, TransRec, and CORE. This instability shows that relational expressiveness is task- and model-dependent. BMS is best interpreted as a structured behavioral-state representation, not as a complete world model: it lacks action-conditioned dynamics, environmental transitions, rewards, causal interventions, and long-horizon simulation.
A broader BWM can combine BMS-like structure with temporal state evolution. A time-indexed behavioral graph can encode current attributes and relations, while a dynamics model predicts how the graph changes under actions, context, and exogenous events. Such an extension would preserve compositional behavioral structure while adding prediction, planning, uncertainty, and counterfactual reasoning.
4. Predictive behavior models for embodied agents
Several BWM formulations focus on embodied behavior rather than symbolic behavioral attributes. The Behavior Foundation Model (BFM) for humanoid robots models reusable distributions over whole-body actions and trajectories across locomotion, teleoperation, motion tracking, and other control modes (Zeng et al., 17 Sep 2025).
BFM treats control modes as alternative goal specifications rather than separate policies. Its inputs include observable proprioception, a unified goal interface, and a binary mask selecting active goal components. The goal interface includes root translation, root orientation, linear and angular velocity, rigid-body link positions, and motor joint angles. The model outputs target joint positions for low-level PD control.
BFM uses a CVAE to model conditional action distributions. Its latent variable is a continuous, state-dependent behavior code. The decoder is conditioned on observable proprioception and the latent code but omits the goal state, encouraging the latent to encode behavior rather than merely copy the current goal. Masked online distillation uses a privileged proxy policy trained in simulation. The BFM is rolled out, its visited states and masked goals are collected, the proxy supplies reference actions, and the BFM is updated through a DAgger-style action loss and KL alignment between privileged and deployable latent distributions.
The model supports latent interpolation and extrapolation. Interpolating root-control and keypoint-control latents produces a combined roundhouse-kick behavior, while latent extrapolation can improve alignment with a desired control mode. A frozen BFM can be adapted to novel behaviors with a residual decoder whose output is added to the pretrained action.
BFM is a behavior foundation model and a goal-conditioned policy prior, but not a complete BWM under a strict predictive definition. It models what behavior should be generated rather than what the environment will do after the action. It does not explicitly predict future proprioceptive states, contacts, terrain evolution, object dynamics, or actuator consequences.
A more predictive embodied model is the Semantic Belief-State World Model (SBWM) for 3D human motion prediction (Chaudhry, 7 Jan 2026). SBWM replaces direct autoregressive pose extrapolation with latent dynamical simulation on the SMPL-X human-body manifold. It maintains a recurrent deterministic belief state 2, a stochastic transition variable 3, and an emission model for SMPL-X parameters.
During observation, the belief state is updated from encoded observations and latent transitions. During open-loop rollout, the model samples 4 from a learned prior, advances 5, and emits future body configurations. The stochasticity is placed in the latent transition rather than independently perturbing each predicted pose, allowing alternative futures to remain temporally coherent.
SMPL-X provides an anatomically structured observation and emission space. The model is intended to separate static body geometry and sensor noise from behaviorally relevant variables such as motion phase, intent, momentum, balance, and contact-related dynamics. These variables are hypotheses induced by the architecture and rollout objective rather than supervised disentanglement results.
On a 15-frame rollout, SBWM reports MPJPE of 61.3, velocity error of 0.048, acceleration error of 0.082, and motion persistence of 0.038. The corresponding deterministic RNN values are 87.2, 0.091, 0.164, and 0.004. In an ablation, removing the stochastic latent produces a 42% freeze rate, while the full latent-plus-belief-feedback model produces a 4% freeze rate. SBWM also improves Best-of-6 MPJPE from 61.3 at 7 to 54.9 at 8.
SBWM is a specialized human-behavior predictive model, not a complete embodied BWM. It lacks external environment state, action conditioning, contact interaction, rewards, and planning. It nevertheless establishes an important BWM pattern: a persistent belief state, stochastic latent transitions, structured body representations, and free-running simulation.
5. Action-conditioned simulation and interactive world models
The Boundless World Model (BWM) is an action-conditioned visual simulator for robot manipulation (Team, 31 Jul 2026). It predicts future camera observations from an initial scene, a dynamic visual history, and temporally aligned robot actions. It adapts the Wan2.2-TI2V-5B video diffusion model with robot-specific action conditioning.
The model combines:
- Initial-environment guidance: a persistent representation of the original scene.
- Dynamic visual history: a moving window of recent observed or generated frames.
- Temporally aligned action conditioning: frame-level cross-attention and latent-level AdaLN conditioning.
- Stateful autoregressive prediction: generated observation chunks are fed back into subsequent context.
The reported configuration uses eight history frames, 72-frame prediction chunks, three prepended boundary actions, an action grouping factor of four, and 14-dimensional absolute end-effector pose commands. The model predicts a distribution over future observations conditioned on the initial frame, current history, and future action chunk.
Its data pipeline uses trajectory replay, overlapping clip sampling, high-resolution rerendering, and initial-observation enhancement. The purpose is not merely to improve visual sharpness but to preserve the temporal correspondence between actions and observations. Rerendering at 480p produces an EWMScore of 63.51 and trajectory accuracy of 64.36, outperforming native 240p, resizing, and generic super-resolution on trajectory-related measures.
BWM serves two roles. As a data engine, it generates action-aligned trajectories for imitation learning. As a policy evaluator, it rolls out candidate policies in closed loop, estimates task outcomes, anticipates risk, and ranks policies before physical execution. On WorldArena, BWM produces 98% success on adjust-bottle and 91% on click-bell, averaging 94.5%, compared with 71.5% for the real-data baseline and 58% for the strongest alternative simulator data source.
For physical robot policy evaluation across six tasks, including folding a towel, opening a drawer, stacking cups, pushing a ball, pushing a block, and wiping a whiteboard, BWM reports a mean absolute error of 14.67 and Pearson correlation 9 when failures are included. Adding BWM-generated data raises mean physical policy success to 71.00%, compared with 50.67% using real data only.
This visual BWM differs from an explicit physics simulator. It does not expose exact object poses, contact forces, or guaranteed physically valid transitions. Its state is implicit in visual context and temporal memory, and generated frames may drift or violate unobserved physical constraints. Its contribution is a domain-specific, action-responsive simulator that can provide useful functional predictions despite lacking a verifiable physical state.
A related control-centered model is GigaBrain-WBC-0.5, which is presented as a BWM for humanoid whole-body control (Cheng et al., 18 Aug 2026). It jointly predicts the next action, next proprioceptive state, and a Gaussian-mixture distribution over the next latent behavior command. A causal Transformer receives proprioception, previous actions, and a reference window, and its auxiliary next-state and next-command objectives force the representation to encode contact dynamics and feasible behavior distributions.
At deployment, the predicted command distribution becomes an admissibility model. Implausible commands are retracted toward a learned Gaussian-mixture component rather than simply rejected. This produces “best-effort” execution: the controller preserves as much of the requested behavior as possible while moving the command toward a learned feasible-behavior region.
In MuJoCo evaluation, GigaBrain-WBC-0.5 reports 81.3% terrain success, 83.1% success under implausible commands, and 99.3% fall recovery. It achieves 96.3% standard flat-ground success and 76.6 mm standard MPKPE. The model is not a general physical simulator or long-horizon planner; its BWM is narrow and action-oriented, modeling proprioceptive consequences and feasible behavior commands for humanoid control.
6. Mental, social, and user-centered behavior models
Physical state alone is insufficient for predicting situated human behavior. Mental World Modeling (MWM) proposes a coupled physical–mental state containing objects, geometry, visibility, environment, beliefs, attention, goals, intentions, emotions, dispositions, norms, constraints, relationships, and social atmosphere (Fei et al., 29 Jul 2026).
The target agent acts from a target-specific partial observation rather than from the global state. This observation includes physically accessible information and mental inferences about other agents. Such inferences may be incorrect, enabling modeling of ignorance, false belief, deception, surprise, and misinterpretation.
MWM decomposes an action into a physical carrier and mental-semantic content. Moving a bag and asking another person to move the bag may produce similar physical access but differ in politeness, ownership, belief correction, and relationship consequences. The transition model therefore updates physical and mental states jointly.
Mentis, the paper’s training-free implementation, parses a global state, generates target-specific observations, decomposes candidate actions, simulates coupled physical and mental successors, and evaluates branches for physical plausibility, mental consistency, and social appropriateness. On Menti-Bench, full structured MWM reaches average final-action F1 of 87.9, compared with 63.3 for direct answering. Removing mental state reduces F1 to 75.8, removing physical state reduces it to 71.4, and decoupling physical and mental transitions reduces it to 81.5. These results support the inclusion of mental variables and their coupling to physical dynamics.
MWM is not intended to simulate consciousness and does not establish human-like Theory of Mind in any model. Its mental variables are task-relevant external representations. The framework remains limited by a small manually constructed benchmark, six-option candidate actions, one-step transitions, LLM hallucination, cultural variation, and reliance partly on LLM judges.
A complementary user-centered perspective is provided by BehaviorBench, which models decision-making from real wallet-level behavioral traces (Yang et al., 1 Jun 2026). The benchmark separates Belief prediction, which predicts final revealed stance and confidence in a market, from Trade prediction, which predicts transaction direction and amount. It contains 141,445 Belief instances and 1,485,972 Trade instances across 2,000 evaluation wallets.
The benchmark compares four history interfaces: no personalization, direct recent history, generated user profiles, and retrieved support-wallet evidence. Profile-based history is generally strongest for Belief prediction, while direct recent history is generally strongest for Trade prediction. This demonstrates that user state is not a single static persona. Stable semantic tendencies and confidence patterns support market-level stance prediction, whereas local sequential and market-specific context supports next-action prediction.
The benchmark reports no-personalization baselines of 53.13% Belief choice accuracy, 0.3417 confidence MAE, 52.85% Trade direction accuracy, and 31.79 Trade amount median absolute error. GPT-5.4 reaches 73.20% Belief choice accuracy with generated profiles and 75.38% Trade direction accuracy with direct recent history. These labels are behavioral proxies, not ground-truth private beliefs or rational preferences.
BehaviorBench also establishes important validity constraints for BWMs. Wallet identity is pseudonymous and does not establish a person’s psychological identity. Public transactions omit private beliefs, off-chain hedges, liquidity constraints, and intent. Observational behavior prediction does not identify causal responses to interventions. Chronological splits, target-wallet-disjoint retrieval, pre-target history, and temporal causality are therefore essential.
7. Evaluation, limitations, and research directions
BWM evaluation must measure behavioral utility rather than rely on one-step reconstruction or latent representation quality. WorldTest proposes reward-free exploration followed by evaluation in a related but different environment (Warrier et al., 22 Oct 2025). Its AutumnBench instantiation contains 43 interactive grid-world environments and 129 tasks spanning masked-frame prediction, planning, and causal change detection.
The protocol separates learning from reward optimization. Agents explore a base environment without external reward, then face a hidden-parameter challenge environment. Evaluation asks whether the acquired model supports action-conditioned prediction, long-horizon planning, and detection of changes to causal dynamics. Humans outperform Claude 4 Sonnet, OpenAI o3, and Gemini 2.5 Pro across all three task families. Additional computation improves performance in 25 of 43 environments but leaves performance flat or worse in 18.
WorldTest also reveals exploration and belief-revision failures. Humans use approximately 12.5% resets and 12.5% no-ops, whereas reasoning models use less than 7% combined resets and no-ops. Models often continue applying an outdated rule after observations contradict it. These findings motivate BWM evaluation that includes information-seeking behavior, uncertainty reduction, counterfactual prediction, and belief revision.
For text-based simulators, Behavior Consistency Reward (BehR) evaluates whether a predicted state induces the same downstream action as the real state under a frozen Reference Agent (Huang et al., 15 Apr 2026). The reward compares normalized per-token log-likelihoods of a logged next action under predicted and real states. It is intended to correct the limitations of Exact Match, token F1, BERTScore, and ROUGE, which may reward preservation of irrelevant text while missing decision-critical errors.
In WebShop, BehR improves pairwise consistency for Qwen3-8B from .345 to .483, for Qwen3-32B from .455 to .485, and for GPT-4o from .760 to .840. Across 16 displayed comparisons, BehR improves 13, ties 3, and degrades none. Its limitations include dependence on the Reference Agent, coverage of only logged actions, possible likelihood miscalibration, offline distribution shift, and the absence of guaranteed long-horizon alignment.
For autonomous-driving BWMs, ReactSim-Bench isolates reactive capability by externally imposing AV trajectories that differ from logged behavior (Zhang et al., 12 Jun 2026). Surrounding-agent simulators must respond to the realized counterfactual AV behavior. The benchmark contains 2,636 scenarios categorized as longitudinal, directional, and lateral deviations. It evaluates collisions, risky time-to-collision, agent–agent collisions, off-road behavior, wrong-way violations, acceleration infeasibility, and steering-curvature infeasibility.
The distinction between realism and reactivity is decisive. Models that closely reproduce logged trajectories can still fail when an independently controlled AV deviates. At 2 Hz replanning, TrajTok reports the lowest agent–AV collision count, 0.1407, while SMART reports the lowest risky agent–AV count, 0.3976. The benchmark demonstrates that log similarity, ADE, and kinematic likelihood do not substitute for interaction-aware closed-loop evaluation.
Across these approaches, recurring BWM requirements are:
- Action-conditioned dynamics: predict how alternative actions change future states and observations.
- Persistent state and memory: retain episodic, semantic, physical, social, and uncertainty-related information.
- Partial observability: distinguish global state from target-specific accessible state.
- Structured behavior representation: preserve relations among attributes, entities, events, and latent behavioral modes.
- Stochasticity: represent coherent alternative futures rather than independent output noise.
- Long-horizon rollout: train and evaluate beyond teacher-forced one-step prediction.
- Reactive interaction: test behavior under externally imposed deviations and changed dynamics.
- Mental and social state: model beliefs, goals, intentions, norms, relationships, and attention where human behavior requires them.
- Calibration and uncertainty: support abstention, safe alternatives, confidence estimation, and uncertainty propagation.
- Counterfactual evaluation: test unseen goals, altered action consequences, object relocation, action remapping, and causal interventions.
- Behavior-based metrics: measure planning success, task preservation, reactivity, policy ranking, calibration, and transfer rather than only reconstruction error.
The principal limitations of current BWM formulations are fragmentation, domain specificity, incomplete state semantics, insufficient formalization, and limited causal guarantees. TM lacks probabilistic semantics; BMS lacks temporal dynamics; the manipulation HMM lacks rich predictive simulation; BFM lacks explicit environmental dynamics; SBWM lacks action and interaction modeling; visual robot BWMs lack verifiable physical state; MWM relies on approximate mental variables; user-behavior benchmarks measure observational rather than causal prediction; and reactive-driving benchmarks focus on one capability among many.
A complete BWM would therefore combine structured state representation, target- or agent-specific observation, action-conditioned stochastic transitions, hierarchical event and behavior representations, predictive emissions, multi-agent interaction, mental or intent variables where appropriate, uncertainty calibration, and behavior-based evaluation. Its defining property would not be a particular neural architecture. It would be the ability to use an internal model of behaviorally relevant dynamics to predict, simulate, evaluate, and select actions under novel goals, partial observability, environmental interaction, altered dynamics, and long-horizon consequences.