Measuring the Value of World-Model Updates: A Counterfactual Utility Protocol for Continual Adaptation
Abstract: Continual world models must decide whether new data justify changing the model. Fixed replay schedules and prediction-error triggers specify when to update, but neither reveals the value of an individual update: one deployment run cannot show how the same model would have performed at that moment had it held its parameters. We introduce the fork ledger, which branches a deployment stream at pre-registered decision points into matched update and hold continuations under common random numbers. It evaluates both continuations on the same episodes and records . Always applying one fixed update mechanism lowers return on all three simulated control tasks: CartPole (; checkpoint-bootstrap CI , against a converged return near $650$), Walker (; ) and Cheetah (; ). Divergence is an outcome of applying the update, so the estimand counts every attempted fork; restricted to the $693$ of $720$ that did not collapse, CartPole and Walker are unchanged in sign ( and ) and Cheetah becomes unresolved (; ). The task is the unit of inference: each contributes $240$ attempted forks over five pretrained checkpoints crossed with two drift directions. The ledger makes counterfactual utility observable for a fixed mechanism, allowing triggers to be judged by the updates they select rather than by surprise detection alone.
Paper Prompts
Sign up for free to create and run prompts on this paper.