Papers
Topics
Authors
Recent
Search
2000 character limit reached

BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling

Published 19 Jun 2026 in cs.RO | (2606.21014v1)

Abstract: Robots must generate trajectories that remain faithful to learned expert behavior while satisfying safety constraints and task-specific objectives specified only at inference time. We formulate constrained trajectory generation for pretrained diffusion and flow-matching policies as Bayesian posterior sampling, with the learned demonstration distribution as a prior and an inference-time, cost-derived likelihood tilting it toward feasible, optimal trajectories. To sample from this posterior without any retraining of the base policy, we leverage the Feynman--Kac corrector framework, originally formulated for diffusion models, and extend it to deterministic flow-matching policies. The result is a unified, inference-time, retraining-free sampler for diffusion and flow policies. We validate the approach on pretrained Diffusion Policy, GR00T-N1.6, and π<em>0.5π<em>{0.5} checkpoints across simulated and real-world manipulation tasks, including planning around non-convex obstacles introduced at inference time, and show improvements over the base π</em>0.5π</em>{0.5} on zero-shot tasks.

Summary

  • The paper introduces a retraining-free Bayesian framework that formalizes constrained robot trajectory generation as posterior sampling, ensuring precise constraint satisfaction.
  • It extends the Feynman-Kac corrector to deterministic flow policies by recasting them as marginal-preserving SDEs, unifying diffusion and flow-based methods.
  • Empirical and theoretical analysis demonstrates exponential concentration of the posterior with near-zero collision rates and high task success in simulated and real-world settings.

Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling

Problem Formulation and Bayesian Perspective

"BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling" (2606.21014) proposes a unifying, Bayesian framework for inference-time constrained robot trajectory generation using pretrained generative policies. Diffusion and flow-matching policies have demonstrated the ability to generate expressive, multimodal behaviors consistent with expert demonstrations. However, trajectory constraints and task objectives—such as collision avoidance, joint limits, or goal-directedness—are generally only known at inference. Existing methods for constrained generation either use post-hoc correction (e.g., gradient-based planners, safety filters), which can yield out-of-distribution or suboptimal samples, or heuristically modify the policy's guidance dynamics without guaranteeing true posterior sampling.

This work formalizes constrained trajectory generation as Bayesian inference: the demonstration distribution learned by the generative policy acts as a prior, while a cost function encoding inference-time constraints is exponentiated to a likelihood, forming a cost-tilted posterior. Sampling from this posterior ensures generated trajectories remain close to the expert prior and are strongly biased towards constraint satisfaction and task-optimality.

Feynman-Kac Sampling and Its Extension to Flow Policies

The key technical contribution is the development of a principled inference-time sampler that targets the exact posterior over trajectories described above without retraining the base policy. The authors leverage the Feynman-Kac (FK) corrector framework, previously introduced for stochastic diffusion models, and extend it to deterministic flow-matching policies:

  • Diffusion Policy Setting: The FKC modifies the standard reverse-time SDE by introducing a drift correction given by the gradient of the cost (i.e., VJ(x)VJ(x)) and accumulates importance weights reflecting the discrepancy between the simulated SDE and the targeted posterior.
  • Flow Policy Extension: By recasting the deterministic flow ODE as a marginal-preserving SDE, using the score function imputed via the learned velocity field, the FKC framework can be applied to deterministic flows, unifying diffusion and flow-guided sampling.

The formalism ensures that, regardless of the underlying generative sampler, the resulting particle system (with importance reweighting) converges to the exact Gibbs-tilted posterior dictated by the prior and cost-induced likelihood. Unlike linear-combination or heuristic guidance, this procedure inherits strong guarantees from the theory of Feynman-Kac functionals.

Theoretical Guarantees

The paper provides a rigorous analysis demonstrating exponential concentration of the posterior: with sufficiently large penalty weights and inverse temperature parameter β\beta, the probability of sampling infeasible or suboptimal trajectories vanishes exponentially fast. Thus, the method eliminates the need for post-hoc safety filters; as β\beta increases, samples concentrate arbitrarily close to the feasibility set and cost minimizers.

Experimental Validation

Extensive empirical studies are conducted on simulated and real-world robotic manipulation tasks across diffusion and flow backbones:

  • On standard manipulation benchmarks (RoboMimic, LIBERO, RoboLab), BayesFP achieves the best or tied-best collision rates and success rates compared to previous state-of-the-art baselines (Linear-Combination Guidance, CASF, JM2D).
  • Notably, BayesFP consistently succeeds in zero-shot constraint satisfaction scenarios, even when complex, non-convex obstacles are introduced only at test time, while maintaining the behavioral diversity imbued by the data-driven prior.
  • In real-world experiments with a SO101 robot, BayesFP is shown to enforce collision avoidance and trajectory mode selection (e.g., suppressing a demonstrated mode at inference) that cannot be achieved via language or observation conditioning alone.

Notable Numerical Results

BayesFP achieves:

  • Zero or near-zero collision rates (e.g., 0–4%) and high task success (up to 100%) in both simulated and real settings, substantially outperforming unconstrained and heuristic-guided baselines across scenarios, especially under non-convex constraints.
  • In real-world bimodal mug selection, the correct-mug placement rate surges from 72% (no guidance) to 100% (BayesFP), with overall success rising from 64% to 92%.

Design Considerations, Ablations, and Limitations

Ablation studies reveal that increasing the number of FK particles improves population coverage and that increasing the inverse temperature β\beta monotonically reduces constraint violations, confirming theoretical predictions. The FKC weight update is essential; omitting it (i.e., using only drift) reduces the approach to linear-combination guidance, which does not recover the correct posterior and underperforms.

The method introduces a modest computational overhead (approximately 1.25–2×\times over vanilla policy sampling), mainly due to the sequential nature of SDF construction and the need for differentiable penalty costs. Implementation is more involved compared to ad hoc guidance techniques due to the need for weight management and resampling. However, parallelization over particles mitigates most practical issues.

Implications and Future Directions

BayesFP establishes a rigorous Bayesian paradigm for incorporating inference-time constraints and task objectives into generative robot policies. The framework is not specific to collision costs; it can, in principle, accommodate any differentiable objective, paving the way for unified, inference-time task composition (e.g., semantic guidance, human correction) via posterior sampling. This allows practitioners to endow generalist policies with new behaviors and safety properties without expanding the training distribution or network size—a crucial property for scalable robotics deployment.

Further, the extension of Feynman-Kac sampling to deterministic flows via marginal-preserving SDEs bridges two major families of generative models. Future directions include applying the framework to low-latency streaming policies, richer constraint representations (e.g., advanced SDFs, learned energy functions), and real-time closed-loop control.

Conclusion

BayesFP (2606.21014) introduces a principled, retraining-free, Bayesian posterior-sampling approach for constrained trajectory generation with pretrained diffusion and flow-based policies. By extending the Feynman-Kac corrector methodology to the unified SDE/ODE setting of modern robot policies, this method guarantees exact posterior sampling with respect to both demonstration data and inference-time constraints. Theoretical and experimental evidence demonstrates superiority over heuristic guidance baselines, particularly in tasks with multimodal solution spaces and non-trivial feasible sets. The framework anticipates a new generation of adaptive, safety-aware, and expressive generative robot policies.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.