Two-Timescale Reinforcement Learning for Real-Time Optimization and Economic NMPC: Experimental Validation
Abstract: We propose a reinforcement learning (RL) framework that tunes both Real-Time Optimization (RTO) and Economic Nonlinear Model Predictive Control (ENMPC) to address plant--model mismatch in process systems. Drawing on modifier-adaptation concepts, the method parameterizes the dynamic model, stage costs, constraints, and RTO modifiers, and uses Q-learning to adjust these parameters at two timescales: a fast update for the ENMPC layer and a slow update for the RTO layer. The framework is experimentally validated on a laboratory rig emulating a three-well subsea oil-production network. Using plant measurement data, the proposed RTO-RLMPC scheme achieves 8.6% higher economic profit than nominal ENMPC, preserves input feasibility of the deployed ENMPC policy and empirically satisfies the path constraints under disturbances and measurement noise, and drives the learned model parameters toward reference plant values, independently identified from experimental data, along the directions that most affect the economic objective. This work provides one of the first experimental demonstrations of RL-tuned ENMPC integrated with RTO.
Paper Prompts
Sign up for free to create and run prompts on this paper.