Robust market making under persistent directional order flow

Establish robust profitability for market-making agents trained in stationary order-flow simulators when exposed to persistent directional flow, overcoming the fragility that is not resolved by simply increasing network size or training-data volume.

Background

The paper identifies robustness under non-stationary order flow as a central unresolved challenge. Stationarily trained market-making agents, including both the GLFT benchmark and the Rainbow distributional DQN controller, can accumulate inventory at their hard limits during persistent buy- or sell-pressure regimes, producing adverse-selection losses and large drawdowns. The proposed Bayesian regime-belief augmentation, queue-adjusted quote-exposure signal, and scenario-bandit fine-tuning substantially improve performance, but the authors continue to characterize robustness as an open problem rather than a fully solved one.

The unresolved scope includes behavior beyond the paper's calibrated zero-intelligence Santa Fe simulator and beyond its fixed regime width, exponentially distributed regime durations, unit-sized quotes, and single-level quoting. The paper's stress tests show that mean profitability can be preserved, but downside tail risk remains under persistent and correlated directional regimes.

References

Third, robustness rather than raw profitability is the central open problem: agents trained in a stationary order-flow simulator are strongly unprofitable under persistent directional flow, and this fragility is not resolved by larger networks or more training data alone.

Deep Learning of Robust Market Making under Regime-Switching Order Flow  (2609.11614 - Moret et al., 10 Sep 2026) in Discussion and Conclusion, immediately before the 'Limitations' paragraph