Papers
Topics
Authors
Recent
Search
2000 character limit reached

Programmable World Model

Updated 15 September 2026
  • Programmable world models are action-conditioned, executable, composable, inspectable, editable, planner-facing, persistent, and verifiable representations of environments that support various interventions and simulations.
  • Use cases include graph-rewriting for chemical retrosynthesis, executable Python simulators for grid-world planning, PDDL environments for task-driven behaviors, and hybrid models combining perceptive and physical dynamics.
  • Effective implementation of a programmable-world-model requires active system identification, program synthesis, interactive verification, and clear separation between observation and intervention semantics

A programmable world model is an executable, inspectable, and configurable model of an environment that represents state, actions, transition dynamics, observations, objectives, constraints, and—where required—persistent rules or mechanisms. Unlike a conventional next-state predictor, it is intended to support intervention, simulation, planning, inspection, modification, and reuse. Its implementation may be symbolic, programmatic, neural, hybrid, or domain-specific. Representative formulations include graph-rewriting world programs, executable Python simulators, products of programmatic experts, symbolic PDDL environments, code-defined web worlds, latent predictive models with explicit interfaces, and hybrid physical simulators.

The concept is not associated with one canonical architecture. Its common denominator is an explicit computational interface through which an agent can encode a state, apply or simulate actions, predict consequences, evaluate futures, and update or modify the model. Depending on the system, this interface may be formalized as W=(A,A(s),T)\mathcal{W}=(\mathcal{A},\mathcal{A}(s),\mathcal{T}), as a code-defined transition St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t), as an executable program F(x)=WMF(x)=WM, or as an action-conditioned latent transition pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a).

1. Conceptual foundations and defining properties

The programmable-world-model perspective extends model-based reinforcement learning by treating the model as more than a predictor of consequences for a supplied action. In ordinary model-based RL, the action space is typically known and the learned model approximates T(s,a,s′)\mathcal{T}(s,a,s'). A programmable model may additionally induce or expose the action vocabulary, determine which actions are applicable, represent action preconditions and effects, and provide an executable mechanism for composing actions.

A formal graph-based formulation defines a world program as:

W=(A,A(s),T),\mathcal{W}=(\mathcal{A},\mathcal{A}(s),\mathcal{T}),

where A\mathcal{A} is the action vocabulary, A(s)\mathcal{A}(s) is the state-dependent action-availability function, and T\mathcal{T} is the transition distribution. In graph-structured domains, actions may be graph-rewriting rules p:L⇝Rp:L\leadsto R, applied by matching a left-hand-side pattern and replacing it with a right-hand-side pattern. The resulting program is executable, reusable across compatible states, and suitable for search or reinforcement learning (Segler, 2019).

A second formulation treats the world model as code synthesized from specifications. WorldCoder represents dynamics and goal-conditioned rewards as Python programs St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)0. The code must reproduce observed transitions and rewards while satisfying an optimism condition requiring an executable trajectory to a positively rewarded terminal state. This makes program synthesis simultaneously explanatory and goal-directed (Tang et al., 2024).

A programmable world model generally has the following properties:

  • Action-conditioned: it represents consequences under interventions rather than only passive continuations.
  • Executable: predictions can be generated by running code, symbolic rules, or a transition mechanism.
  • Composable: states, actions, objects, rules, and plans can be combined into larger structures.
  • Inspectable: internal state, transition logic, action schemas, and planning traces can be examined.
  • Editable: rules, parameters, modules, constraints, or code can be changed without necessarily retraining the complete system.
  • Planner-facing: imagined trajectories can be evaluated by search, MPC, MCTS, value iteration, or another decision procedure.
  • Persistent: state and consequences can survive beyond the current observation or frame.
  • Verifiable: execution, invariants, tests, replay, and failure traces can be used to assess behavioral correctness.

The distinction between prediction and simulation is central. A visual generator may produce plausible frames without exposing state, action semantics, alternative interventions, or planning interfaces. An actionable world model instead provides a sandbox for hypothetical reasoning: given a state and candidate action, it generates possible futures that can be compared against goals, rewards, or constraints (Xing et al., 7 Jul 2025).

The channel being modeled is also significant. A formal analysis distinguishes the environment channel St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)1, the agent channel St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)2, and the realized joint process St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)3. These channels induce different predictive states. A model learned from on-policy trajectories may predict the joint agent–environment process without supporting arbitrary action interventions. Programmability therefore requires declaring whether the model is intended for environment simulation, agent prediction, or closed-loop interaction modeling (Baltieri et al., 23 Jul 2026).

2. Representational regimes

Programmable world models employ several representational regimes, each exposing different forms of structure and control.

Structured symbolic and graph representations

Graph-based models represent states as multisets of graphs, with vertices as entities and edges as relations. This supports variable-size, compositional environments such as molecules, source code, text, images represented as lattice graphs, and physical systems. Graph-rewriting programs provide explicit action semantics and permit symbolic search (Segler, 2019).

WorldPrediction evaluates a related capability at a higher semantic level. Its WorldPrediction-WM task asks a model to select the high-level action that explains an initial-to-final visual transition, while WorldPrediction-PP asks it to select an ordered sequence of temporally extended actions. The benchmark treats actions as options with initiation conditions, low-level policies, and termination conditions, and tests hidden intermediate-state inference and temporal composition. Current frontier models achieve 57.0% on WM and 38.1% on PP in the reported results, compared with effectively perfect human performance after filtering (Chen et al., 4 Jun 2025).

Executable programs and simulators

WorldCoder represents a deterministic, fully observed, episodic contextual MDP through Python transition and reward programs. Object-centric symbolic states contain entities, types, coordinates, and attributes. The same transition program can be reused across goals by replacing or modifying the reward program. WorldCoder uses code execution for imagined trajectories and a conventional planner for action selection, rather than querying the LLM at every action (Tang et al., 2024).

Agent2World generalizes this approach to PDDL domains and executable Python environments. Its generated artifacts encode predicates, actions, preconditions, effects, state transitions, rewards, termination conditions, and simulator APIs. The system emphasizes that syntactic validity is insufficient: a model must be executed, unit-tested, and evaluated through interactive simulation (Hu et al., 26 Dec 2025).

Web World Models use ordinary web code, TypeScript interfaces, APIs, procedural generators, schemas, caches, and databases as the world substrate. The state is divided into a deterministic Physics Layer St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)4 and a model-generated Imagination Layer St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)5:

St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)6

The code layer determines authoritative state transitions, while an LLM generates narratives, descriptions, dialogue, and other context. This separation provides controllability, persistence, graceful degradation, and reproducibility conditional on stable seeds, prompts, model versions, and caching (Feng et al., 29 Dec 2025).

Modular probabilistic programs

PoE-World represents a world model as an exponentially weighted product of programmatic experts:

St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)7

Each expert expresses a local mechanism, such as horizontal motion, vertical motion, object deletion, contact, or death after enemy collision. Experts may leave attributes unspecified, allowing other experts to provide those predictions. Weights are fitted by maximum likelihood; low-weight experts can be pruned, and hard constraints can exclude physically implausible outcomes. This structure provides modularity, stochasticity, partial specification, and localized online revision (2505.10819).

Latent state representations

Latent models provide compact internal states for prediction and rollout. World Machine introduces a continuous latent world state St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)8 and trains a transformer to predict sensory data through that state. Its State Discovery procedure repeatedly feeds predicted states back into training, while masking, sequence breaking, local attention, temporal recall, and bounded St+1ϕ=fcode(Stϕ,at)S_{t+1}^{\phi}=f_{\mathrm{code}}(S_t^{\phi},a_t)9 activations encourage latent-state stability. Prediction from a single stored state remains difficult, but the architecture demonstrates state-based continuation and variable-context inference (Nascimento et al., 21 May 2026).

WorldModelLens provides a capability-typed interface across recurrent latent models, token-based models, and joint-embedding predictors. Its required operations are encode, transition, initial_state, and sample_z; optional capabilities include decoding, reward prediction, continuation prediction, actor execution, and critic evaluation. This creates a common execution and intervention substrate without requiring all models to share the same internal architecture (Challagundla et al., 7 Jun 2026).

The limitations of purely latent representations are substantial. Latent variables may be opaque, non-identifiable, insufficiently grounded, or unsafe to edit. A latent predictor may also model an on-policy joint process rather than an intervention-capable environment. Consequently, latent programmability requires explicit action semantics, state introspection, uncertainty handling, and validation.

3. Construction, learning, and verification

Programmable-world-model construction ranges from induction from transition examples to LLM-based code synthesis, active system identification, and hybrid physics modeling.

Action and transition induction

In graph-based world programs, state-state examples are used to identify changed and invariant graph components. Localized edits are converted into reusable graph-rewriting actions. The learned action vocabulary may be very large: approximately 300,000 reaction rules were induced in the chemical retrosynthesis application. A neural action-prior model estimates F(x)=WMF(x)=WM0, retaining rules whose cumulative probability exceeds 99%, and a transition classifier assesses whether proposed successor states are compatible with learned dynamics (Segler, 2019).

Program synthesis with optimism and debugging

WorldCoder begins with random exploration and a replay buffer containing F(x)=WMF(x)=WM1. An LLM synthesizes or revises Python transition and reward programs. Data consistency requires exact reproduction of observed next states, rewards, and termination flags. Optimism requires the program to entail at least one executable path from the initial state to a positively rewarded terminal state.

Program refinement is treated as a bandit problem. Candidate programs are arms, refinement success is binary reward, and Thompson sampling selects programs for further refinement. Execution failures, incorrect predictions, and runtime errors are returned as debugging feedback. The resulting system transfers transition code across environments and often changes only the reward program for new natural-language goals (Tang et al., 2024).

Agent2World makes verification a central part of world-model generation. Its three-stage architecture consists of a Deep Researcher, a Model Developer, and an adaptive Testing Team. The Researcher fills specification gaps through web search. The Developer creates PDDL or executable simulator artifacts. The Testing Team performs generated unit tests and interactive simulation. Successful multi-turn trajectories become verifier-guided supervision for fine-tuning.

On Text2World, Agent2World Multi achieves 93.1% executability and 75.4 average F1 for GPT-4.1-mini. On the Code World Models Benchmark, it obtains overall accuracy 0.5441 and normalized return 0.4811. The fine-tuned Llama-3.1-8B system improves overall CWMB accuracy from 0.3150 to 0.3754 and normalized return from 0.2296 to 0.3197. These results support behavioral verification as a complement to static parsing and textual similarity (Hu et al., 26 Dec 2025).

Qualitative structure selection and parameter fitting

VisualPatchWorld explicitly separates qualitative dynamics from continuous parameters. Short active probes select a hypothesis F(x)=WMF(x)=WM2 from an environment-specific family F(x)=WMF(x)=WM3, after which a Python template F(x)=WMF(x)=WM4 is fitted using multi-step prediction error. The induced program may represent joint-space kinematics, contact-driven pushing, grip-gated motion, or linear navigation.

Its central transition contract is:

F(x)=WMF(x)=WM5

where F(x)=WMF(x)=WM6 is a structured scene graph. The program is rolled forward during CEM-MPC, and selected action prefixes are executed before replanning. In matched executable-code comparisons, VPW achieves 69.0% mean planning success, compared with 45.5% for the strongest code baseline. Hybrid validation with a physics engine raises PushT success from 22% to 96% with oracle state (Bai et al., 28 Jul 2026).

Compilation and typed interfaces

World Craft introduces a semantic intermediate representation F(x)=WMF(x)=WM7 between natural-language intent and executable scene layout:

F(x)=WMF(x)=WM8

The final scene F(x)=WMF(x)=WM9 contains metadata, assets, layout, and properties such as collision, navigation, semantic tags, agents, and persistent character references. An Enricher creates pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)0, a Manager grounds it into layout, a Critic performs validation and repair, and an Artist retrieves or synthesizes visually consistent assets (Sun et al., 14 Jan 2026).

ChemWorld applies the same principle to chemical experimentation. A compatibility compiler verifies dependencies, units, state ownership, resource requirements, instruments, and lifecycle closure before constructing a world. Typed actions are executed transactionally through preflight validation, runtime checks, candidate execution, post-execution validation, commit or rollback, observation generation, and receipt recording (Qiu et al., 11 Aug 2026).

Hybrid physical and scientific models

Domain-specific systems may combine learned perception with analytic or executable dynamics. MetaEI-WM constructs a geometry-, semantic-, material-, and radio-aware electromagnetic world model from robot sensing, semantic parsing, material assignment, and ray tracing. It predicts radio propagation under candidate information-metasurface states and computes coding configurations through physics-informed inverse design. The resulting loop connects natural-language targets, digital-twin construction, electromagnetic simulation, physical actuation, RSSI or CSI feedback, and local refinement (Liu et al., 2 Jul 2026).

ODesign applies a related idea to biomolecular interaction design. It uses shared modality-aware tokens, all-atom diffusion, Pairformer interaction representations, inverse folding, hotspot conditioning, rigid and flexible target modes, and partial diffusion. Its programmability is structural rather than temporal: users specify modalities, fixed or redesignable entities, epitopes, motifs, atoms, lengths, and diffusion centers (Zhang et al., 25 Oct 2025).

4. Planning, execution, and interaction

A programmable world model becomes operational when connected to a planner or controller. Several planning regimes are used.

Search and reinforcement learning

The graph-based world program constructs an internal simulator pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)1. It supports imagined rollouts, MCTS, best-first search, heuristic planning, and model-based RL. In chemical retrosynthesis, PUCT-MCTS achieves a 95.24% solve rate at 13.0 seconds per plan on 497 target molecules, outperforming the other reported search methods in the combined success-time comparison (Segler, 2019).

WorldCoder uses depth-limited value iteration over executable Python dynamics and rewards. It concentrates LLM computation in synthesis and debugging, while ordinary program execution supports repeated planning. In the reported MiniGrid-UnlockPickup setting, the optimistic system learns a useful model in no more than approximately 100 actions, whereas PPO fails after pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)2 actions under the reported comparison (Tang et al., 2024).

Model-predictive control

VPW uses CEM-MPC. Candidate action sequences are rolled forward through the induced program, scored by goal distance, and partially executed before replanning. Hybrid scoring evaluates all candidates with the induced program and checks a selected top fraction in the ground-truth engine. This reduces physics queries while preserving planning performance in contact-rich domains (Bai et al., 28 Jul 2026).

SWM supplies a standardized environment and planning substrate. Its World abstraction supports synchronous environments, attached policies, dataset recording, evaluation, factor-of-variation control, and MPC solvers including CEM, MPPI, SGD, and Adam. The system makes environment configuration, data collection, planner selection, horizon, receding horizon, and evaluation programmable, although it does not permit arbitrary new physical laws without extending the environment implementation (Maes et al., 9 Feb 2026).

Hierarchical and abstract planning

PoE-World uses an abstract graph for Montezuma’s Revenge. Nodes represent contact-based symbolic states, and edges represent simulated reachability. Breadth-first search finds a high-level route, while low-level MCTS or greedy action chunks execute subgoals. Constraints mainly function as damage control for long-horizon planning: they increase success from pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)3 to pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)4 on Montezuma’s Revenge and from pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)5 to pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)6 on its altered variant (2505.10819).

Transactional interaction and replay

ChemWorld treats every action as an auditable transaction. Failed actions can be rejected before execution or rolled back after candidate execution. Physical state and observation randomness are restored on rollback, while declared attempt costs may remain in the resource ledger. The complete action trace—including rejected and rolled-back operations—can be replayed exactly under the same software identities, seeds, world identity, intervention specification, and submitted trace (Qiu et al., 11 Aug 2026).

Interactive agent–environment coupling

A model that predicts only realized trajectories may be compact but unsuitable for arbitrary interventions. The distinction between unrestricted environment models and support-restricted models is therefore operationally important. A support-restricted model can be derived from a joint model by treating unsupported actions as transitions to an absorbing failure symbol pf(s^′∣s^,a)p_f(\hat{s}'\mid\hat{s},a)7. This improves compactness but limits counterfactual coverage to action continuations supported by the coupling (Baltieri et al., 23 Jul 2026).

5. Applications and domain-specific instantiations

Programmable world models have been demonstrated or proposed across structured chemistry, games, web environments, electromagnetic systems, biomolecular design, time-series prediction, visual control, and scientific experimentation.

Domain World representation Programmable interface Reported use
Chemical retrosynthesis Molecular graphs and graph-rewrite rules Reaction-rule vocabulary, applicability model, transition model Retrosynthetic planning (Segler, 2019)
Grid-world tasks Python object-centric simulator Transition and reward code, goal-conditioned programs Model-based planning and transfer (Tang et al., 2024)
Atari Product of programmatic experts Weighted local stochastic rules and constraints Pong and Montezuma’s Revenge planning (2505.10819)
Web environments Typed web state and code-defined physics APIs, schemas, procedural generation, persistence Travel, narrative, encyclopedic, and game worlds (Feng et al., 29 Dec 2025)
Biomolecular design Multimodal tokens and all-atom coordinates Hotspots, motifs, atoms, modality, diffusion conditions Protein, nucleic-acid, and ligand interaction design (Zhang et al., 25 Oct 2025)
Visual control Scene graphs and executable dynamics Qualitative sketches, fitted parameters, CEM-MPC Navigation, reaching, pushing, manipulation (Bai et al., 28 Jul 2026)
Electromagnetics Semantic 3-D digital twin with materials IMS coding states and target regions Wireless enhancement and physiological sensing (Liu et al., 2 Jul 2026)
Chemical experimentation Compiled process modules and private laws Typed operations, instruments, resources, world forks Controlled agent experimentation (Qiu et al., 11 Aug 2026)
World creation Semantic layouts and executable scenes Scene schemas, assets, collision, navigation, NPC state AI Town-like visual environments (Sun et al., 14 Jan 2026)
Research infrastructure Configurable benchmark environments Factors of variation, policies, datasets, MPC Reproducible world-model research (Maes et al., 9 Feb 2026)

In chemistry, graph-rewrite programs support recursive decomposition of target molecules into purchasable building blocks. Expert organic chemists preferred PUCT-MCTS routes over heuristic BFS routes at 68.2% versus 31.8%, while the authors caution that this does not establish general reliability across domains (Segler, 2019).

In games and grid worlds, executable code provides persistent rules, object interactions, reward predicates, and transfer across environments. WorldCoder’s code-level transfer preserves movement and turning dynamics while extending object types such as keys and doors. Web World Models broaden this pattern to persistent, procedurally generated environments whose structural state is deterministic while LLMs provide narratives and high-dimensional content (Tang et al., 2024, Feng et al., 29 Dec 2025).

In visual control, VPW demonstrates that explicit qualitative mechanisms can substantially improve planning where joint kinematics or contact dynamics are essential. Its remaining gap on contact-rich pushing illustrates that source-level programmability does not automatically imply accurate physical simulation (Bai et al., 28 Jul 2026).

In scientific domains, metaEI-WM and ChemWorld illustrate two distinct forms of programmability. MetaEI-WM programs physical electromagnetic boundaries through metasurface coding, using an explicit environment model and ray-traced propagation. ChemWorld programs the world itself: researchers change topology, operating conditions, instruments, hidden laws, or resource policies while preserving the public experimental contract (Liu et al., 2 Jul 2026, Qiu et al., 11 Aug 2026).

6. Evaluation, limitations, and research directions

Evaluation must distinguish predictive fidelity, executable correctness, planning utility, controllability, transfer, interpretability, and safety. Low reconstruction error or visual plausibility is insufficient. A programmable world model should be tested under interventions, novel combinations, out-of-distribution states, execution failures, changed goals, and altered rules.

Evaluation dimensions

Important dimensions include:

  • Transition fidelity: whether predicted successors agree with observed or ground-truth transitions.
  • Long-horizon rollout quality: whether small local errors compound during simulation.
  • Planning utility: whether the model improves goal attainment, sample efficiency, or search efficiency.
  • Action coverage: whether the model supports actions outside the behavior-policy support.
  • Controllability: whether changing an action, goal, rule, or constraint produces the intended change.
  • Compositional generalization: whether known entities and mechanisms work in unseen combinations.
  • Execution validity: whether generated programs parse, run, satisfy interfaces, and produce coherent trajectories.
  • Replayability: whether trajectories can be reconstructed under controlled software and randomness conditions.
  • Interpretability: whether state variables, rules, active modules, and causal traces can be inspected.
  • Uncertainty calibration: whether the model identifies unsupported, ambiguous, or unreliable predictions.

WorldPrediction shows that current systems remain weak on high-level causal transitions and long-horizon procedural composition: the best reported results are 57.0% on single high-level transitions and 38.1% on ordered action sequences, compared with effectively perfect human performance after filtering (Chen et al., 4 Jun 2025).

Principal failure modes

Model exploitation and distribution shift are pervasive. Planners may find trajectories that exploit inaccuracies in the learned simulator. States generated by the planner may lie outside the training distribution, causing compounding errors. This problem is reported in graph-based chemical planning, WorldCoder, PoE-World, latent world models, and learned visual dynamics systems (Segler, 2019, Tang et al., 2024, 2505.10819).

Representation errors arise when the chosen state representation omits relevant variables or merges distinct states. GNN embeddings may conflate graph structures; symbolic scene graphs may lack joint configuration; latent vectors may not expose semantically meaningful coordinates; visual perception may recover object identities without accurate metric geometry (Segler, 2019, Bai et al., 28 Jul 2026, Nascimento et al., 21 May 2026).

Action ambiguity and insufficient observability complicate induction from state-state pairs. Multiple graph rewrites may explain the same transition, hidden variables may be omitted, and action labels may be unavailable. In causal modeling, a model learned from on-policy trajectories may not answer unsupported intervention queries (Segler, 2019, Baltieri et al., 23 Jul 2026).

Scalability remains difficult. Large rule libraries cause branching-factor and graph-matching problems; executable programs can become large; LLM-based synthesis and verification are computationally expensive; hybrid physical models may require substantial sensing and GPU resources; and exhaustive long-horizon simulation remains costly (Segler, 2019, Hu et al., 26 Dec 2025, Liu et al., 2 Jul 2026).

Limited stochasticity and partial observability affect many systems. WorldCoder assumes deterministic, fully observed environments. VPW primarily induces deterministic programs. World Machine does not implement action-conditioned control in its experiments. ODesign models static structural generation rather than biological temporal dynamics. PoE-World provides probabilistic experts but relies on approximate factorization and object-centric perception (Tang et al., 2024, Nascimento et al., 21 May 2026, Zhang et al., 25 Oct 2025, 2505.10819).

Verification incompleteness is unavoidable. Unit tests, simulations, replay, and invariant checks can expose many errors but cannot prove correctness over all states and action sequences. Agent2World’s testing team may miss untested counterexamples; ChemWorld’s guarantees apply only to declared modules and bound software identities; Web World Models’ schemas prevent malformed objects but not semantic or factual errors (Hu et al., 26 Dec 2025, Qiu et al., 11 Aug 2026, Feng et al., 29 Dec 2025).

Research directions

A stronger programmable world model would combine several existing ideas:

  1. Typed action and state interfaces from graph programs, WorldModelLens, Web World Models, and ChemWorld.
  2. Executable transition programs from WorldCoder, Agent2World, VPW, and web-based world models.
  3. Probabilistic modularity from PoE-World.
  4. Latent compression and variable-context rollout from World Machine.
  5. Active qualitative system identification from VPW.
  6. Physics and domain priors from metaEI-WM and ODesign.
  7. Behavior-aware verification and replay from Agent2World and ChemWorld.
  8. Controlled environment variation from SWM.
  9. Channel-aware predictive semantics from the computational-mechanics formulation.
  10. Hierarchical planning and hybrid validation from PoE-World, VPW, and graph-based MCTS systems.

The resulting system would need explicit distinctions between observations and interventions, unrestricted and support-restricted predictions, learned and user-defined rules, factual and counterfactual rollouts, and executable state versus generated context. It should expose provenance, uncertainty, action coverage, constraints, version identities, and failure behavior.

The central unresolved issue is the relationship between flexibility and verifiability. Fully learned models offer broad interpolation and perceptual richness but obscure their mechanisms. Fully hand-authored simulators offer inspectability but require substantial construction effort and may not scale. Programmable world models occupy an intermediate regime: they seek to induce or synthesize executable structure from data while retaining enough explicit semantics for planning, testing, editing, and intervention.

A mature programmable world model would therefore be neither merely a video generator nor merely a fixed simulator. It would be a persistent computational substrate in which states, actions, dynamics, goals, constraints, tools, and planners are explicit enough to inspect and modify, yet sufficiently adaptive to learn from observations and generalize to new configurations.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Programmable World Model.