ThermoRL: RL for Thermal Optimization
- ThermoRL is a research program defining reinforcement learning methods that optimize thermal transport, temperature control, and thermostability in diverse systems.
- It integrates techniques from fluid control, energy management, quantum state preparation, and molecular design to achieve quantifiable improvements over traditional methods.
- Key challenges include addressing partial observability, ensuring reward fidelity, and enforcing safety constraints while enhancing sample efficiency and sim-to-real transfer.
Searching arXiv for papers relevant to "ThermoRL" and nearby RL-for-thermal-control usage. ThermoRL is a label used across several research lines for reinforcement-learning formulations in which the optimized object is thermal transport, thermodynamic dissipation, a temperature parameter, or thermostability. In the thermo-fluid literature, it denotes active control of convection, heat transfer, and radiative exchange; in RL methodology, it can denote explicit optimization of sampling temperature or curriculum temperature; in quantum and statistical physics, it refers to entropy-production minimization and thermal-state preparation; and in molecular design, it denotes structure-aware mutation search for enhanced protein thermostability (Beintema et al., 2020, Dang et al., 12 Feb 2026, Wang et al., 24 Jul 2025).
1. Terminological scope and principal meanings
Across the cited literature, “ThermoRL” is not a single standardized algorithm but a family of problem formulations in which reinforcement learning is coupled to thermal or thermodynamic structure. In one usage, the “thermo” component is literal: buoyancy-driven convection, pulsating cooling jets, radiative heat transfer, battery thermal management, hot-water systems, or thermal power plants are controlled by policies that act on boundary temperature, flow actuation, device setpoints, or plant operating points. In another usage, temperature itself becomes the optimized variable, either as a learnable meta-policy in large-language-model RL or as a thermodynamic control coordinate in maximum-entropy curricula. A third usage places RL inside statistical mechanics and quantum thermodynamics, where the objective is to minimize entropy production, discover high-efficiency cycles, or prepare Gibbs, generalized Gibbs, or SYK thermal states. A fourth usage shifts from thermal physics to thermostability, using RL to optimize mutation choices in proteins (Rudolf et al., 2024, Adamczyk et al., 12 Mar 2026, Baba et al., 2022, Wang et al., 24 Jul 2025).
A concise way to organize the term is by the quantity being optimized. In thermo-fluid control the target is typically a transport observable such as the Nusselt number or a thermo-mechanical residual. In energy systems it is an operational cost, line-loading margin, temperature-tracking error, or combustion efficiency. In temperature-centered RL it is the exploration–exploitation balance induced by a sampling temperature. In quantum thermodynamics it is free energy, entropy production, or local-equilibrium error. In protein design it is a surrogate-predicted , with stabilizing mutations defined by (Beintema et al., 2020, Zhan et al., 2021, Dang et al., 12 Feb 2026, Xu, 2021, Wang et al., 24 Jul 2025).
2. Thermo-fluid control and heat-transfer optimization
The most direct ThermoRL usage appears in active control of buoyancy-driven and convective flows. In two-dimensional Rayleigh–Bénard convection under the Boussinesq approximation, a model-free PPO controller observes temperature and both velocity components on an probe grid over the current and previous control steps and actuates lower-wall temperature segments with binary levels , . The reward is , the objective is minimization of time-averaged Nusselt number, and the controller stabilizes the conductive regime up to , compared with the classical threshold and linear-controller stabilization only up to 0. For 1 and up to 2, the learned policy reduces heat flux by a factor of about 3 relative to uncontrolled flow and discovers a double-cell configuration with 4 (Beintema et al., 2020).
A later framework on the same control class attributes earlier degenerate actuation to two deficiencies: MLP policies that discard spatial structure and memoryless policies that cannot separate self-induced flow changes from background evolution. By replacing MLPs with convolutional encoders, adding GRU memory, switching to off-policy TD3 or MADDPG, and enforcing action-smoothness constraints, the method achieves cell coalescence in Rayleigh–Bénard convection at 5 and reduces 6 to as low as 7, 8 below the uncontrolled baseline, within 350 episodes. In double-diffusive salt-finger convection, the same framework discovers a travelling-wave actuation whose phase speed adapts to the evolving mixing state, enhancing heat transfer by 9 and reducing salinity variance by 0 (Cavallazzi et al., 4 Jun 2026).
ThermoRL has also been used for forced-convection control in CFD-coupled settings. In pulsating impinging jets, DQN variants regulate the inlet velocity of a cooling jet over a laminar regime 1. The reward is shaped around maintaining the average surface temperature within 2 of the setpoint 3. Soft Double DQN and Dueling DQN maintain the temperature in the desired threshold for more than 4 of the control cycle, whereas hard Double DQN fails the task (Salavatidezfouli et al., 2023).
In thermal-radiation design, RL is coupled not to a flow solver but to a fluctuational-electrodynamics forward model. For near-field radiative heat transfer between multilayer hyperbolic metamaterials, the state is a binary vector specifying metal or dielectric layers, the action edits the layer stack, and the reward is the heat-transfer coefficient relative to a periodic baseline. Double DQN, PPO, and related methods navigate a discrete design space of size 5, and for 6 the best-found stack reaches 7, a 8 improvement over the physically intuitive periodic design (Ortiz-Mansilla et al., 2024).
In laser additive manufacturing, the emphasis shifts from direct control to reward and world-model diagnosis. A bilevel Proxy–FEA framework evaluates scan-order policies with cheap thermo-inspired proxies at the lower level and sparse Abaqus FEA labels at the upper level. On a simplified LDED32 stripe benchmark, the study finds an observed stress–distortion trade-off rather than a single monotonic quality objective; within the evaluated set, center_out is the robust compromise candidate, while raster_left_to_right and edge_in occupy opposing endpoints. The diagnostic conclusion is that current cheap path-based metrics predominantly capture distortion-related behaviour and exhibit only weak correlation with sparse FEA reference labels, so proxy-only rewards are misaligned for future RL training (Wu et al., 24 May 2026).
3. Thermal management, power systems, and industrial energy applications
In engineered thermal systems, ThermoRL often denotes RL-assisted control or parametrization under embedded safety logic. For battery-electric-vehicle thermal management, the RL agent does not actuate the plant directly; it tunes lookup-table parameters of a gain-scheduled PI valve controller. The observation joins a context vector, an image-like representation of 9 parameter maps for 0 and 1, and ECU signal histories; the action is a continuous masked update of those parameter maps; and DroQ is used as the learning algorithm. In simulation and real-vehicle testing, the learned parameters outperform the supplier baseline in all MAE/RMSE entries across scenarios 1 and 2, while the hybrid “ours + expert” configuration gives the best MAE/RMSE on the dynamic GP-track scenario 3 (Rudolf et al., 2024).
For domestic hot-water systems, a model-based RL framework learns vessel thermodynamics, heater behaviour, and user demand online from minimal sensors. The control objective is energy reduction under strict comfort and safety constraints, with a risk-averse planner selecting reheat decisions subject to hot-water availability and anti-legionella cycles. In a 32-house deployment in the Netherlands, the method reduces energy consumption for DHW production by roughly 2 with no loss of occupant comfort, corresponding to approximately 3 annual savings per household (Kazmi et al., 2018).
A more lightweight thermostat formulation dispenses with explicit thermodynamic simulation and instead encodes empirical human rules of thumb into the reward. The state is outdoor temperature and humidity, the action is a continuous setpoint in 4, and DDPG optimizes a comfort–energy proxy. Evaluated by the area between outdoor and indoor temperature curves, the learned policy reduces the energy proxy from 5 to 6, or about 7, relative to a fixed 8 setpoint (Rao, 2019).
In power-system applications, “thermal” refers to line current limits and thermal plant performance rather than fluid temperature alone. In thermal cascading prevention, an A3C topology controller for the IEEE 14-bus grid uses a physics-based reward 9 with a terminal penalty of 0, together with a curriculum that progressively tightens overload enforcement. On 150 hardest-enforcement scenarios, the curriculum agent successfully reaches the end in 1 cases, whereas the non-curriculum baseline usually terminates within 500 steps (Matavalam et al., 2021). In thermal power generation, DeepThermal formulates combustion optimization as an offline constrained MDP, learns a combustion-process simulator, and trains MORE, a model-based offline RL algorithm with restrictive exploration. Deployed in four coal-fired plants in China, it reports efficiency gains such as 2 at 3, together with reductions in NO4 and exhaust temperature (Zhan et al., 2021).
Related load-control work focuses on thermostatically controlled loads under sparse sensing. One line uses sequence models within fitted Q-iteration; in residential heating and electric water heaters, LSTM-based policies outperform CNN and MLP feature extractors under partial observability, reducing cost by 5 over no control in the heating case and by 6 for the water heater (Ruelens et al., 2017). Another line couples Q-learning to a Modelica–Python co-simulation pipeline for voltage control of TCL populations; with stochastic initialization and constant 7 p.u., the best constant-gain baseline achieves mean MSE 8, while smart discretization and RL provide comparable or slightly better tracking, particularly beyond the training interval (Lukianykhin et al., 2020).
4. Temperature as an explicit RL control variable
A distinct ThermoRL usage treats temperature not as part of the physical plant but as the control variable of the RL procedure itself. In large-language-model RL, Temperature Adaptive Meta Policy Optimization recasts sampling temperature as a categorical meta-policy 9 over candidates 0. The inner loop updates the language-model policy, for example with GRPO, using trajectories sampled at temperature 1, while the outer loop reweights candidate temperatures according to virtual likelihoods and trajectory advantages, without extra rollouts. On five mathematical reasoning benchmarks, TAMPO achieves average accuracy 2 in Pass@1/Pass@8, compared with 3 for the best fixed-temperature baseline at 4, and the learned meta-policy typically favours higher 5 after warmup before gradually shifting lower later in training (Dang et al., 12 Feb 2026).
A more theoretical formulation connects curriculum learning in maximum-entropy RL to non-equilibrium thermodynamics. Reward parameters are interpreted as coordinates 6 on a task manifold, excess work is approximated by 7, and optimal curricula are geodesics of the induced metric. For temperature annealing, the resulting MEW schedule uses the scalar friction 8, estimated from reward autocovariance, and updates the inverse temperature according to a constant-thermodynamic-speed rule. In Humanoid-v5, MEW yields higher and more stable returns than constant-temperature baselines and a SAC-style auto-temperature baseline, while in a 9 Grid World the geodesic curriculum detours around a high-friction barrier (Adamczyk et al., 12 Mar 2026).
These two lines share the same conceptual move: temperature is promoted from a fixed hyperparameter to an optimizable policy variable. The difference is that TAMPO is an online, advantage-guided meta-policy over discrete decoding temperatures, whereas MEW derives an annealing schedule from a thermodynamic geometry of tasks (Dang et al., 12 Feb 2026, Adamczyk et al., 12 Mar 2026).
5. Thermodynamic, statistical-mechanical, and quantum formulations
In statistical physics and thermodynamics, ThermoRL often means that the reward is a path-extensive thermodynamic functional. One example uses neural-network RL to discover high-efficiency heat-engine cycles from elementary reversible processes. An evolutionary algorithm recovers the maximally efficient Carnot cycle, learns the Stirling cycle when adiabats are removed, and finds the Otto cycle under the appropriate action restrictions; when an additional irreversible process is added, it discovers a previously unknown hybrid cycle (Beeler et al., 2019).
Another example studies finite-time shortcuts between equilibrium states of open systems at fixed bath temperature. The control parameters 0 enter either a classical two-level master equation or a quantum GKLS equation, and the reward combines the negative instantaneous entropy-production rate with a terminal penalty enforcing arrival at the target Gibbs state. In the classical case, the RL protocol reproduces the analytically optimal trajectory and the finite-time reachability bound 1; in the quantum two-level case, it finds two-dimensional controls 2 that minimize total entropy production under the same fixed-time constraint (Xu, 2021).
Thermal-state preparation in many-body quantum systems provides a further ThermoRL interpretation. One framework uses recurrent deep Q-learning to prepare pure states whose local observables match Gibbs or generalized Gibbs ensembles. The reward is 3, where 4 collects local observables. For non-integrable systems, preparation errors decay exponentially with system size, consistent with canonical typicality; for integrable systems and GGE targets, the decay is polynomial, consistent with finite-size fluctuations of generalized ensembles (Baba et al., 2022).
In the SYK model, RL is used for circuit-architecture search in thermal-state preparation on quantum hardware. A DDQN with a 3D-CNN encodes the circuit tensor, while a composite reward based on free energy and, in a refined version, fidelity, guides the design of a two-stage PQC. For 5 Majorana fermions, the discovered circuits reduce CNOT counts by two orders of magnitude relative to first-order Trotterization; for instance, at 6 and 7, the reported improvement factors are 8 for the free-energy reward and 9 when fidelity is added (Kundu, 20 Jan 2025).
At the nanoscale, the thermodynamic obstacle becomes the learning signal itself. In thermally dominated control, if 0 is the action-to-thermal ratio, the optimal improvement scales as 1 whereas the learned improvement of model-free RL scales as 2, so learning efficiency vanishes as 3. The proposed remedy is to learn policies at lower temperature, where 4, and transfer them to the operational regime (Boccardo et al., 2023).
6. Thermostability design in proteins
A comparatively new usage shifts ThermoRL from heat and dissipation to molecular stability. In protein mutation design, ThermoRL combines a pre-trained structure-aware GNN encoder, hierarchical Q-learning, and a surrogate 5 predictor to choose both where and which amino-acid substitution to apply. The action is decomposed into a position decision 6 and a residue decision 7, reducing the action space from 8 to 9. The reward is the surrogate-predicted 0, with stabilizing mutations defined as 1. Across training and unseen proteins, the framework achieves higher or comparable cumulative rewards relative to BO-GP, BO-ENN, and random search, filters out destabilizing mutations, and accurately detects key mutation sites in unseen proteins (Wang et al., 24 Jul 2025).
This usage broadens the semantics of “thermo.” The optimized object is not heat transfer or a temperature schedule, but thermostability of a macromolecule measured through free-energy change upon mutation. Even so, the underlying ThermoRL pattern remains recognizable: a sequential policy searches a large constrained design space under a thermodynamic objective supplied by a surrogate model (Wang et al., 24 Jul 2025).
7. Recurring methodological themes and open problems
Several themes recur across otherwise disparate ThermoRL formulations. The first is partial observability and delayed actuation. In Rayleigh–Bénard control, controllability degrades when observation or actuation delays become comparable with the Lyapunov time, and the main bottleneck at high 2 is plausibly actuation delay from the boundary to the bulk (Beintema et al., 2020). In domestic TCL control under sparse observations, LSTM-based sequence models outperform CNN and MLP alternatives because the latent thermal state must be reconstructed from long histories rather than instantaneous measurements (Ruelens et al., 2017).
The second theme is fidelity of rewards and world models. The scan-order study for additive manufacturing shows that cheap proxies predominantly capture distortion-related behaviour and correlate only weakly with high-fidelity FEA labels, implying that proxy-only reward design can misdirect policy search (Wu et al., 24 May 2026). Comparable issues arise in industrial thermal management, where sim-to-real transfer, parameter clipping, action masking, and conservative reward shaping are used because embedded controllers and plant models cannot be trusted to represent the full operational envelope (Rudolf et al., 2024).
A third theme is safety and constraint handling. In hot-water control, backup policies enforce comfort and anti-legionella constraints (Kazmi et al., 2018). In combustion optimization, MORE enforces a CMDP constraint through a cost critic and Lagrangian relaxation while filtering model-generated samples by sensitivity and density (Zhan et al., 2021). In thermo-fluid control, zero-mean projection, amplitude caps, and smoothness penalties regularize wall-temperature actuation into physically plausible patterns rather than saturated or pseudo-random outputs (Cavallazzi et al., 4 Jun 2026).
A plausible synthesis is that ThermoRL is best understood not as a single method but as a research program: reinforcement learning is coupled to thermal physics, thermodynamic functionals, or temperature variables in ways that preserve domain structure rather than treating the environment as a generic black box. The main unresolved issues are therefore also domain-structured: sample efficiency when each rollout requires high-fidelity simulation, principled reward construction when proxies are only partially aligned with target physics, recurrent or delay-aware control under sparse sensing, and transfer from simulation or surrogate supervision to experiments, hardware, or field deployment (Beintema et al., 2020, Wu et al., 24 May 2026, Zhan et al., 2021).