PO4AO: Policy Optimization for Adaptive Optics
- PO4AO is a model-based reinforcement learning framework that optimizes adaptive optics control by leveraging dynamic system modeling and policy prediction.
- It overcomes challenges such as temporal delay errors, mis-registration, nonlinearity, and photon noise using neural network architectures.
- Validated in simulation, laboratory, and on-sky tests, PO4AO consistently outperforms standard integrators by enhancing residual error reduction, Strehl ratios, and contrast.
Searching arXiv for PO4AO and closely related adaptive optics control papers to ground the article in current literature. PO4AO, short for Policy Optimization for Adaptive Optics, is a model-based reinforcement learning framework for adaptive optics (AO) control developed to address limitations of conventional static matrix-based wavefront reconstruction and integrator control, especially temporal delay errors, mis-registration, nonlinearity, photon noise, and rapidly varying observing conditions (Nousiainen et al., 2022). It frames AO control as a Markov Decision Process in which a learned dynamics model predicts the AO system response and a learned policy model outputs deformable-mirror updates from recent wavefront-sensor telemetry and command history (Nousiainen et al., 2022). Across numerical simulations, laboratory experiments on GHOST and MagAO-X, and an on-sky deployment on the Papyrus AO system at the OHP 1.52 m telescope, PO4AO was reported to improve residual error, Strehl ratio, and coronagraphic contrast relative to a standard integrator, while operating in a turnkey manner once tuned for a given system (Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026).
1. Origins and problem setting
PO4AO emerged from work on AO control for high-contrast imaging and, in particular, the direct imaging of exoplanets at small angular separations from their host stars (Nousiainen et al., 2022). In that setting, residual stellar light left by imperfect AO correction propagates into the coronagraphic point spread function and limits detectability. The motivating claim in the PO4AO literature is that current controllers based on static matrix-based wavefront reconstruction and integrator control are robust and computationally efficient, but they remain limited by temporal delay error, dynamic misregistration, calibration or model nonlinearities, and poor handling of vibrations and fast changes in conditions (Nousiainen et al., 2022, Nousiainen et al., 9 Jun 2026).
The method was first studied through numerical simulations of eXtreme Adaptive Optics with Pyramid wavefront sensing for both 8-m and 40-m apertures, then implemented in laboratory environments, and subsequently demonstrated on sky (Nousiainen et al., 2022, Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). This progression establishes PO4AO not as a purely theoretical control law but as a framework tested across increasingly realistic AO settings.
A recurrent theme in the literature is that PO4AO is intended as an automated or turnkey approach to AO control. In the 2022 study, reinforcement learning is described as enabling an AO controller whose usage is “entirely a turnkey operation,” while the 2026 on-sky study states that, once tuned for Papyrus, PO4AO operated in a turnkey fashion, using a single set of hyperparameters across varying observing conditions and science targets (Nousiainen et al., 2022, Nousiainen et al., 9 Jun 2026).
2. Reinforcement-learning formulation and controller structure
PO4AO casts AO control as a Markov Decision Process in which the state is formed from a temporal history of reconstructed wavefront-sensor observations and previous corrective actions, the action is a differential deformable-mirror command, and the reward is tied to the residual wavefront after correction (Nousiainen et al., 2022, Nousiainen et al., 2023).
In one formulation, the state is written as
with denoting the DM-space projection of the wavefront-sensor measurements (Nousiainen et al., 2023). A closely related formulation describes the policy as acting on recent WFS reconstructions and residual DM commands,
where is the learned policy network (Nousiainen et al., 9 Jun 2026).
The dynamics model predicts the next observation from the current state and action,
and the reward or cost proxy is given by the negative squared Euclidean norm of the predicted residual,
(Nousiainen et al., 9 Jun 2026). In the GHOST implementation, the reward is described as the negative squared residual plus a regularization term on the action,
(Nousiainen et al., 2023). This suggests that the precise reward may vary slightly between implementations, while preserving the same basic control objective: minimization of the residual wavefront in the controlled modal space.
Policy optimization is carried out over a planning horizon using the learned dynamics model:
(Nousiainen et al., 9 Jun 2026). The AO command update is then applied through a leaky-integrator-like recursion,
where 0 is the leak factor (Nousiainen et al., 9 Jun 2026). For comparison, the standard integrator is written as
1
(Nousiainen et al., 9 Jun 2026).
Both the policy and the dynamics model are described as fully convolutional neural networks, typically shallow, with the AO system’s spatial grid structure used for computational efficiency (Nousiainen et al., 2022, Nousiainen et al., 2023). The 2022 study further states that the policy output is projected onto a physically motivated Karhunen-Loève modal basis, filtering uncontrollable or poorly sensed modes (Nousiainen et al., 2022). In the 2026 on-sky analysis, a linear approximation to the learned controller showed that PO4AO distributes control non-locally and temporally, unlike an integrator, which is diagonal and instantaneous (Nousiainen et al., 9 Jun 2026). A plausible implication is that the controller’s performance derives not only from temporal prediction but also from learned cross-modal structure in the control law.
3. Training procedure, hyperparameters, and real-time implementation
A central operational feature of PO4AO is that training and inference are performed in parallel, which is repeatedly identified as crucial for eventual on-sky use (Nousiainen et al., 2023). The procedure begins with a warm-up phase in which data are collected using a standard integrator plus added noise or random exploration, so that the initial policy and dynamics networks can be trained on a sufficiently broad set of states (Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026).
The papers describe a two-thread or dual-path implementation. In the GHOST test bench, a Python module contained a control thread that reads WFS-processed voltages from shared memory, runs policy inference, writes DM commands, and logs transitions, plus a training thread that updates the policy and dynamics models continuously in parallel (Nousiainen et al., 2023). The Papyrus implementation at OHP similarly interfaced a Python controller with the existing DAO RTC via shared-memory buffers, enabling deployment without replacing the core real-time controller (Nousiainen et al., 9 Jun 2026). The 2026 study explicitly describes the approach as compatible with real AO hardware and implementable non-disruptively as a user-space module (Nousiainen et al., 9 Jun 2026).
The literature identifies several important hyperparameters. On Papyrus, the tuned values included episode length of 1000 steps, approximately 2 s at 500 Hz, history length of 64 frames, number of controlled modes equal to 209, and planning horizon 2 (Nousiainen et al., 9 Jun 2026). The GHOST paper states that performance is especially sensitive to the history length and planning horizon, which must align with the system delay and dominant disturbance periods (Nousiainen et al., 2023). This is consistent with the role of PO4AO as a predictive controller: too short a history or too short a simulated horizon would undercut its ability to learn temporal structure.
Real-time feasibility is a recurring engineering issue. The 2022 study reports inference time of roughly 0.3 ms per step for an 8-m simulation and 0.35–0.7 ms for the ELT-scale case on GPU hardware (Nousiainen et al., 2022). The GHOST implementation introduced approximately 700 microseconds of additional latency beyond hardware, pipeline, and Python interface latency (Nousiainen et al., 2023). In the on-sky Papyrus deployment, the Python implementation introduced approximately 720 microseconds mean latency with 3 microseconds and up to 5% frame drops, while the paper abstract summarizes this as approximately 4 of additional latency, plus control jitter and occasional frame drops (Nousiainen et al., 9 Jun 2026). The authors nonetheless report performance gains despite these prototype-level software constraints, and they identify lower-level implementations such as C/C++, CUDA, or TensorRT as paths toward lower latency (Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026).
4. Simulation and laboratory validation
The first detailed evaluations of PO4AO were conducted in numerical simulation and laboratory experiments rather than on sky. The 2022 study used COMPASS simulations for a VLT-scale system with 1364 DM actuators and a 5 PWFS, and an ELT-scale system with over 10,000 actuators and a 6 PWFS (Nousiainen et al., 2022). In those simulations, PO4AO was reported to improve post-coronagraphic contrast by factors of 4–7 over the integrator for a 0-mag star with a non-modulated PWFS, by 3–9 under noisier 9 mag conditions, and by 20–90 in an ideal WFS case (Nousiainen et al., 2022). The same work states that PO4AO could be trained on timescales of 5–10 seconds and that inference remained below 1 ms, which was presented as sufficient for real-time XAO control (Nousiainen et al., 2022).
Laboratory demonstrations followed on MagAO-X and later GHOST. The 2022 paper reports that on the MagAO-X laboratory setup, PO4AO reduced speckle variance in science images by factors of 3–7 at 7–8 relative to the integrator (Nousiainen et al., 2022). The 2023/2024 GHOST study emphasized different strengths: transparent adaptation to increasing, unknown control-loop delays of 0, 1, or 2 full frames, strong low-flux performance down to S/N 9, and robustness under dynamic misregistration with the DM shifted by up to 40% of actuator spacing (Nousiainen et al., 2023). In that work, PO4AO outperformed the optimally tuned integrator by a factor 0 in wavefront residuals in the time-delay experiments and remained stable without recalibration after system changes that destabilized a classical integrator unless recalibrated (Nousiainen et al., 2023).
These studies also provide a mechanistic interpretation of PO4AO’s gains. The method is described as predictive, self-calibrating, and capable of handling non-linear wavefront sensing, including the Pyramid WFS optical gain effect (Nousiainen et al., 2022, Nousiainen et al., 2023). The GHOST experiments showed that longer history improves control performance and enables damping of periodic disturbances, including a 16 Hz vibration, which supports the claim that PO4AO uses temporal windows not just for noise averaging but for learned predictive rejection of structured disturbances (Nousiainen et al., 2023).
5. On-sky demonstration on Papyrus
The first on-sky validation of PO4AO was reported in 2026 and constitutes the first on-sky demonstration of a reinforcement learning controller for adaptive optics (Nousiainen et al., 9 Jun 2026). The controller was deployed on the Papyrus adaptive optics system installed at the Coudé focus of the 1.52 m telescope (T152) at the OHP. Papyrus uses a pyramid WFS, a 241-actuator DM, an IR science path, and a loop frequency of 500 Hz (Nousiainen et al., 9 Jun 2026).
The on-sky experiments compared PO4AO against a standard integrator controller over several nights and across different guide-star magnitudes, flux levels, seeing conditions, and the presence of telescope vibrations (Nousiainen et al., 9 Jun 2026). The study reports that PO4AO consistently outperformed the standard integrator in all tested configurations (Nousiainen et al., 9 Jun 2026). Strehl ratios extracted in the paper are summarized below.
| Target | PO4AO Strehl (%) | Integrator Strehl (%) |
|---|---|---|
| Vega | 12.1 | 10.5 |
| Vega (2nd night) | 26.5 | 18.2 |
| HD 177809 | 22.45 | 12.86 |
| Cygni | 18.9 | 9.9 |
| Alpha Cas | 13.3 | 8.0 |
The paper also reports relative peak intensity ratios of 0.77, 0.74, 0.68, 0.60, and 0.62 for the integrator relative to PO4AO in those configurations, again favoring PO4AO (Nousiainen et al., 9 Jun 2026). Under a night affected by a malfunctioning tracking motor, PO4AO reduced the tilt-mode vibration PSD peak by a factor 4.5 at 80 Hz, whereas the integrator could not (Nousiainen et al., 9 Jun 2026). Residual variance for most modes was lower by factors of 1.2–2.3, with the largest gains on low-order modes (Nousiainen et al., 9 Jun 2026). Under poor seeing, reported as 2.2″ to 3.8″ in general and up to 4″ on the third night, PO4AO still delivered noticeably higher Strehl and peak intensity (Nousiainen et al., 9 Jun 2026).
A notable operational conclusion from the on-sky study is that once the controller had been tuned for Papyrus, it used the same hyperparameters for all tested targets and conditions (Nousiainen et al., 9 Jun 2026). This is the most concrete basis in the literature for the description of PO4AO as a turnkey AO controller.
6. Distinctive properties, comparisons, and limitations
The PO4AO papers consistently position the method against the classical integrator, and, more broadly, against Kalman/LQG and model-free RL approaches (Nousiainen et al., 2023). The stated advantages include predictive capability, self-calibration, robustness to misregistration, tolerance of photon noise, and operation under nonlinear sensing conditions (Nousiainen et al., 2022, Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). In the comparison table given in the GHOST study, PO4AO is characterized as nonlinear, data-driven, robust, able to self-calibrate, able to use arbitrary history via neural networks, and fast in online adaptation relative to classical alternatives (Nousiainen et al., 2023).
Several features recur as technically distinctive.
First, PO4AO learns both control and reconstruction from data, rather than relying on a fixed interaction matrix and manually tuned gains (Nousiainen et al., 2023). Second, it performs continuous adaptation through online learning, allowing it to track changes in wind, optical gain, misregistration, and noise level (Nousiainen et al., 2022, Nousiainen et al., 2023). Third, it uses temporal history explicitly, which enables learned compensation for latency and periodic disturbances (Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). Fourth, it can act as a drop-in replacement within existing AO data flow, since the real-time controller still provides WFS preprocessing and command transport while PO4AO supplies the control law (Nousiainen et al., 2022, Nousiainen et al., 2023).
At the same time, the papers do not present PO4AO as free of engineering constraints. All three major studies emphasize latency sources introduced by the machine-learning implementation, especially Python overhead, CPU-GPU synchronization, and contention between inference and training threads (Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). The on-sky deployment succeeded despite additional latency and frame drops, but the paper explicitly states that performance was achieved in spite of a non-optimized Python implementation (Nousiainen et al., 9 Jun 2026). This indicates that the observed performance should not be conflated with an upper bound on the method; rather, the authors argue that optimization could further improve results.
A common misconception is that reinforcement learning controllers for AO had remained purely simulated or laboratory-based. That characterization became outdated with the Papyrus result, which the 2026 paper presents as the first on-sky demonstration of such a controller (Nousiainen et al., 9 Jun 2026). Another misconception is that PO4AO requires per-target manual retuning; the on-sky study reports the opposite for the tested Papyrus configurations (Nousiainen et al., 9 Jun 2026). A more defensible caution is that turnkey operation was demonstrated after tuning for a specific system, not as a guarantee of zero commissioning effort across arbitrary AO architectures.
7. Prospects and scope
The current literature presents PO4AO as a controller designed for single-conjugate AO and XAO use cases, with a clear emphasis on high-contrast exoplanet imaging (Nousiainen et al., 2022, Nousiainen et al., 9 Jun 2026). Its demonstrated scope ranges from laboratory systems such as MagAO-X and GHOST to on-sky operation on Papyrus (Nousiainen et al., 2022, Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). The 2026 paper proposes broader deployment to more challenging and higher-order systems, including MagAO-X, SCExAO, and upgraded VLT/SPHERE/SAXO+, and identifies extension to more nonlinear wavefront sensing scenarios, such as Zernike and non-modulated pyramid WFS, as future directions (Nousiainen et al., 9 Jun 2026).
The method is especially relevant to regimes where existing AO controllers encounter known failure modes: delay-dominated control, non-stationary calibration, vibration contamination, low-flux operation, and dynamic misregistration (Nousiainen et al., 2022, Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). The 2022 simulations further suggest that the approach scales to ELT-size systems, with sub-millisecond inference reported for configurations exceeding 10,000 actuators (Nousiainen et al., 2022). The 2026 paper adds that future work should examine retraining frequency, parallelization strategies, and real-time determinism for systems with 5,000–10,000+ actuators (Nousiainen et al., 9 Jun 2026).
Taken together, the published results define PO4AO as a model-based RL controller that learns a temporal-spatial control policy directly from AO telemetry, uses an internal dynamics model to optimize that policy over a short horizon, and has now been validated from simulation through laboratory operation to on-sky deployment (Nousiainen et al., 2022, Nousiainen et al., 2023, Nousiainen et al., 9 Jun 2026). The available evidence supports its characterization as a predictive, self-calibrating, and system-compatible AO control framework, while also indicating that software optimization remains a significant part of its maturation path.