Papers
Topics
Authors
Recent
Search
2000 character limit reached

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

Published 6 Jul 2026 in cs.LG | (2607.05252v1)

Abstract: Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. However, existing methods struggle with effective multimodal modeling. They often rely on brute-force fusion strategies that ignore the structural disparities between parameters and observations, thus limiting estimation fidelity. In this work, we introduce FUSE (Feynman-Kac steered mUlti-modal flow matching for efficient Simulation-based posterior Estimation). Unlike prior work, FUSE employs a dual-track architecture that preserves the distinct features of multimodal inputs while facilitating dynamic interaction. Additionally, we propose an FK-steered sampling strategy that leverages intermediate observation likelihoods to guide the generative trajectories, effectively improving the sample quality during inference. Our approach outperforms state-of-the-art baselines on standard SBI benchmarks, producing posteriors that closely match ground-truth MCMC. Furthermore, in a real-world exoplanet orbital estimation task, FUSE successfully resolves complex parameter degeneracies that challenge existing methods, highlighting its potential to accelerate complex scientific discoveries in astrophysics and beyond.

Summary

  • The paper introduces a dual-track MM-DiT architecture with FK-steered flow matching to enhance simulation-based inference precision and data efficiency.
  • It employs a conditional rectified objective and likelihood-guided FK-steering for adaptive resampling, improving posterior fidelity.
  • Empirical results show FUSE reduces simulation time and outperforms state-of-the-art methods on complex benchmarks including exoplanet inference.

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

Introduction

FUSE presents a simulation-based inference (SBI) framework tailored for high-fidelity, efficient posterior estimation in complex scientific settings characterized by heterogeneous, multi-modal data. The methodology addresses two critical limitations of prevailing approaches to generative posterior modeling in SBI: suboptimal multimodal data fusion and ineffective utilization of simulator-based likelihood information during amortized inference. By introducing a dual-track, Feynman-Kac (FK)-steered flow-matching backbone—anchored in an MM-DiT architecture and augmented with a test-time likelihood-guided resampling mechanism—FUSE establishes new data-efficiency and fidelity benchmarks, especially for tasks exhibiting significant parameter degeneracy and sharp multimodality.

Problem Background and Limitations of Prior Work

Scientific applications such as exoplanet orbital inference and gravitational wave characterization demand posterior estimates that capture not only high SNR structure but also subtle multimodalities and parameter covariances. Standard neural SBI methods (NPE, FMPE, Simformer) either aggregate context and parameters via brute-force vector concatenation or rely on static summary embeddings. These strategies are agnostic to the physical structure of the data, failing to preserve modality-specific characteristics and leading to underspecified posteriors—particularly in the presence of degenerate mappings and heterodimensional conditional relationships between data and parameters. Conventional amortized neural samplers are computationally efficient but can yield physically implausible samples. Asymptotically exact MCMC approaches remain intractable for large-scale or urgent inference due to prohibitively high per-sample simulation costs.

Methodology

Multimodal Diffusion Transformer (MM-DiT) Backbone

FUSE employs a dual-track MM-DiT architecture to simultaneously retain modality-specific representations and facilitate bidirectional interaction:

  • Tokenization: Physical parameters and observations are independently embedded via modality-specific projections, mapping parameters to multiple tokens per dimension for enhanced representational capacity.
  • Fusion: Token sequences undergo joint multimodal self-attention, allowing context-dependent parameter refinement. Layerwise iterative message passing disentangles cross-modal dependencies, enabling the backbone to retain salient modality structure across the generative trajectory.
  • Posterior Prediction: The backbone's parameter sub-sequences are globally aggregated via shared, lightweight linear heads, allowing for efficient, permutation-invariant mapping from token space to the flow-matching velocity domain.

Flow Matching with Conditional Rectified Objective

The generative process is formulated as a conditional flow-matching ODE on parameter space, learned via regression of the velocity field to the optimal transport path connecting base noise and data samples. The objective employs a conditional rectified-flow loss, enforcing the learned vector field to approximate the target transport between latent and parameter configurations at any intermediate time.

Feynman-Kac (FK) Steered Test-Time Correction

Offline-trained amortized samplers are augmented at inference by an FK-steering mechanism:

  • Particle Propagation: Multiple trajectories are propagated in parallel using the learned flow-matching dynamics.
  • Simulator-Based Reward: At selected intermediate steps, the predicted (denoised) parameters are scored using the true simulator-based log-likelihood and prior, forming a reward signal for each trajectory.
  • Likelihood-Guided Resampling: Based on these scores, particle populations are adaptively resampled to allocate computational resources toward high-posterior-density regions, correcting for amplitude misallocation caused by transport approximation errors.
  • Stochastic Rejuvenation: SDE-based noise perturbations are injected post-resampling, preserving sample diversity and mitigating mode collapse due to resampling.

The FK framework thus seamlessly integrates physical likelihood information into amortized sampling, enhancing high-density coverage with marginal increased inference overhead relative to the base amortized sampler.

Empirical Results

Synthetic SBI Benchmarks

Empirical evaluation on 10 canonical SBIBM tasks demonstrates quantitatively superior posterior fidelity compared to NPE, FMPE, and Simformer, as measured by â„“-C2ST, MMD, KL, Sinkhorn distance, PME, PVR, and median error metrics. Notably, FUSE achieves the lowest average â„“-C2ST (0.59) and KL (0.28), establishing state-of-the-art alignment with reference MCMC posteriors, especially in simulation-rich regimes. FK-steering is ablated during benchmark runs, isolating the architectural contribution.

  • Sample Efficiency: As simulation budget scales, FUSE outperforms baselines on tasks with strong degeneracy (e.g., LV, SLCP), while the MM-DiT backbone proves robust and data-efficient.
  • Ablation: Use of individual parameter tokenization and intermediate-layer fusion is critical for modeling complex posterior topology.

FK-Steered Inference on High-Difficulty Tasks

On the SLCP benchmark and exoplanet inference for β Pictoris b (8D), inclusion of FK-steering further refines sample concentration, boosting high-density region coverage (posterior mode accuracy improves, credible region IoU increases). FK-steering yields posteriors that closely match the dominant MCMC reference structures and outperform both static best-of-N selection and standalone amoritized backbones, especially in parameter subspaces exhibiting strong degeneracy (e.g., orbital inclination, longitude of ascending node).

  • Efficiency: FUSE completes high-precision exoplanet inference in minutes, whereas PTMCMC requires hours on multi-core infrastructures.
  • Robustness: FK-steering's resampling and stochasticity avoid under-dispersion despite non-asymptotic particle counts in test-time correction.

Theoretical Implications

FUSE demonstrates that the integration of FK-steered flow matching within a multimodal-transformer backbone can realize amortized SBI with posterior fidelity previously unobtainable outside full MCMC pipelines. Theoretically, FK-steering implements a pathwise posterior tilting in the Feynman-Kac measure, leveraging the learned reverse-time sampler as proposal. Likelihood evaluation at the denoised mean ensures physical validity of simulator calls, and stochastic rejuvenation preserves local exploration.

Limitations concern the absence of formal MCMC-type asymptotic guarantees—particularly with finite-particle FK steering and in regimes of severe simulation paucity, where overfitting or poor tail coverage can manifest. Furthermore, while scalability has been demonstrated up to the exoplanet benchmark, performance in higher-dimensional, real-time settings (e.g., gravitational wave parameter inference) remains subject to empirical validation.

Practical and Theoretical Impact

Practically, FUSE enables orders-of-magnitude reduction in turnaround time for complex astrophysical inference tasks, making real-time updating and intervention plausible as new data become available. The modularity of MM-DiT and FK-steering supports extensibility to other scientific inverse problems and multimodal generative modeling tasks in physical sciences. Theoretically, this work bridges advances in diffusion-based multimodal generative modeling and adaptive likelihood-guided correction, providing a clear demonstration that amortized neural architectures can approach high-fidelity posterior coverage without direct MCMC sampling.

Open directions include rigorous control for tail coverage in the finite-particle regime, architectural refinement for extreme data scarcity, and extension to sequential or online SBI with dynamic adaptation.

Conclusion

FUSE establishes a new paradigm for simulation-based posterior estimation in inverse scientific problems, leveraging a dual-path multimodal transformer with FK-steered likelihood guidance to achieve both rapid, amortized, and highly structured posterior inference. The combination outperforms existing neural and classical SBI methods in both fidelity and efficiency, enabling practical deployment in time-critical scientific pipelines. While finite-particle FK correction remains an approximation, the approach validates the effectiveness and generalizability of likelihood-steered amortized generative samplers for multimodal, heterogeneous inference, with immediate application to astronomy and broader scientific discovery.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 2 likes about this paper.