Extend joint distribution flow matching with additional training views and reinforcement learning

Develop extensions of joint distribution flow matching for synthesizer inversion that incorporate additional views of the training data for improved off-manifold robustness and use reinforcement learning to refine the learned distribution for off-manifold inputs.

Background

The paper introduces joint distribution flow matching as a way to model audio and synthesizer parameters jointly, enabling training with paired synthetic data and unpaired real recordings. Experiments show improved inversion robustness for real-world audio that lies off the synthesizer's audio manifold.

The conclusion identifies unresolved extensions enabled by the framework: incorporating further views of the training data may improve robustness and provide additional conditioning and control mechanisms, while reinforcement-learning fine-tuning may further refine the learned distribution for off-manifold inputs. These possibilities are presented as open directions rather than results established by the paper.

References

Several directions remain open. The flexibility of the joint distribution flow matching framework enables us to include further ``views'' of our training data which may further improve robustness, and unlocks further options for conditioning and control. Moreover, our model's probabilistic framing suits it well to fine-tuning by reinforcement learning, which may serve to further refine the learnt distribution for off-manifold inputs.

— Off-manifold robustness in synthesizer inversion with joint distribution flow matching  (2609.29320 - Hayes, 24 Sep 2026) in Section 6, Conclusion