---
title: 'PO4AO: Policy Optimization for Adaptive Optics'
url: https://www.emergentmind.com/topics/po4ao
type: topic
---

# PO4AO: Policy Optimization for Adaptive Optics

Searching arXiv for PO4AO and closely related adaptive optics control papers to ground the article in current literature.
PO4AO, short for **Policy Optimization for Adaptive Optics**, is a **model-based reinforcement learning** framework for adaptive optics (AO) control developed to address limitations of conventional static matrix-based wavefront reconstruction and integrator control, especially **temporal delay errors**, **mis-registration**, **nonlinearity**, **photon noise**, and rapidly varying observing conditions [2205.07554]. It frames AO control as a Markov Decision Process in which a learned **dynamics model** predicts the AO system response and a learned **policy model** outputs deformable-mirror updates from recent wavefront-sensor telemetry and command history [2205.07554]. Across numerical simulations, laboratory experiments on GHOST and MagAO-X, and an on-sky deployment on the Papyrus AO system at the OHP 1.52 m telescope, PO4AO was reported to improve residual error, Strehl ratio, and coronagraphic contrast relative to a standard integrator, while operating in a **turnkey** manner once tuned for a given system [2401.00242] [2606.10771].

## 1. Origins and problem setting

PO4AO emerged from work on AO control for **high-contrast imaging** and, in particular, the direct imaging of exoplanets at small angular separations from their host stars [2205.07554]. In that setting, residual stellar light left by imperfect AO correction propagates into the coronagraphic point spread function and limits detectability. The motivating claim in the PO4AO literature is that current controllers based on **static matrix-based wavefront reconstruction** and **integrator control** are robust and computationally efficient, but they remain limited by **temporal delay error**, **dynamic misregistration**, **calibration or model nonlinearities**, and poor handling of **vibrations** and fast changes in conditions [2205.07554] [2606.10771].

The method was first studied through numerical simulations of **eXtreme Adaptive Optics** with **Pyramid wavefront sensing** for both **8-m** and **40-m** apertures, then implemented in laboratory environments, and subsequently demonstrated on sky [2205.07554] [2401.00242] [2606.10771]. This progression establishes PO4AO not as a purely theoretical control law but as a framework tested across increasingly realistic AO settings.

A recurrent theme in the literature is that PO4AO is intended as an **automated** or **turnkey** approach to AO control. In the 2022 study, reinforcement learning is described as enabling an AO controller whose usage is “entirely a turnkey operation,” while the 2026 on-sky study states that, once tuned for Papyrus, PO4AO operated in a **turnkey fashion**, using a **single set of hyperparameters across varying observing conditions and science targets** [2205.07554] [2606.10771].

## 2. Reinforcement-learning formulation and controller structure

PO4AO casts AO control as a **Markov Decision Process** in which the **state** is formed from a temporal history of reconstructed wavefront-sensor observations and previous corrective actions, the **action** is a differential deformable-mirror command, and the **reward** is tied to the residual wavefront after correction [2205.07554] [2401.00242].

In one formulation, the state is written as
$$
\bm{s}_t = \begin{pmatrix} \bm{o}_{t}, \bm{o}_{t-1}, \dots, \bm{o}_{t-k}, \bm{a}_{t-1}, \dots, \bm{a}_{t-k} \end{pmatrix},
$$
with $\bm{o}_t = C_m \Delta \bm{w}_t$ denoting the DM-space projection of the wavefront-sensor measurements [2401.00242]. A closely related formulation describes the policy as acting on recent WFS reconstructions and residual DM commands,
$$
\bm a_t = \pi_\theta(\bm s_t) = \pi_\theta(\bm o_t, ..., \bm o_{t-k}, \bm a_{t-1}, ..., \bm a_{t-k}),
$$
where $\pi_\theta$ is the learned policy network [2606.10771].

The **dynamics model** predicts the next observation from the current state and action,
$$
\tilde{\bm o}_{t+1} = \hat p_\omega(\bm s_t, \bm a_t),
$$
and the reward or cost proxy is given by the negative squared Euclidean norm of the predicted residual,
$$
\hat r_\omega(s, a) = - \|\tilde{\bm o}_{t+1}\|^2
$$
[2606.10771]. In the GHOST implementation, the reward is described as the negative squared residual plus a regularization term on the action,
$$
\hat r_\omega (s, \bm s_{t+1}, a) = - \|\bm{o}_{t+1}\|^2 - \alpha \|a\|^2
$$
[2401.00242]. This suggests that the precise reward may vary slightly between implementations, while preserving the same basic control objective: minimization of the residual wavefront in the controlled modal space.

Policy optimization is carried out over a planning horizon $H$ using the learned dynamics model:
$$
\arg\max_\theta \sum_{\bf s} \in \mathcal{D} \sum_{t=1}^{H} \hat r_\omega(\tilde s, \pi_\theta(\tilde s))
$$
[2606.10771]. The AO command update is then applied through a **leaky-integrator-like** recursion,
$$
\bm v_t = \alpha \bm v_{t-1} + \bm a_t,
$$
where $\alpha$ is the leak factor [2606.10771]. For comparison, the standard integrator is written as
$$
\bm{v}_t = \alpha \bm{v}_{t-1} + g C_m \Delta \bm{w}_t
$$
[2606.10771].

Both the **policy** and the **dynamics model** are described as **fully convolutional neural networks**, typically shallow, with the AO system’s spatial grid structure used for computational efficiency [2205.07554] [2401.00242]. The 2022 study further states that the policy output is projected onto a physically motivated **Karhunen-Loève modal basis**, filtering uncontrollable or poorly sensed modes [2205.07554]. In the 2026 on-sky analysis, a linear approximation to the learned controller showed that PO4AO distributes control **non-locally and temporally**, unlike an integrator, which is diagonal and instantaneous [2606.10771]. A plausible implication is that the controller’s performance derives not only from temporal prediction but also from learned cross-modal structure in the control law.

## 3. Training procedure, hyperparameters, and real-time implementation

A central operational feature of PO4AO is that **training and inference are performed in parallel**, which is repeatedly identified as crucial for eventual on-sky use [2401.00242]. The procedure begins with a **warm-up phase** in which data are collected using a standard integrator plus added noise or random exploration, so that the initial policy and dynamics networks can be trained on a sufficiently broad set of states [2401.00242] [2606.10771].

The papers describe a two-thread or dual-path implementation. In the GHOST test bench, a Python module contained a **control thread** that reads WFS-processed voltages from shared memory, runs policy inference, writes DM commands, and logs transitions, plus a **training thread** that updates the policy and dynamics models continuously in parallel [2401.00242]. The Papyrus implementation at OHP similarly interfaced a Python controller with the existing **DAO RTC** via **shared-memory buffers**, enabling deployment without replacing the core real-time controller [2606.10771]. The 2026 study explicitly describes the approach as compatible with real AO hardware and implementable non-disruptively as a **user-space module** [2606.10771].

The literature identifies several important hyperparameters. On Papyrus, the tuned values included **episode length** of **1000 steps**, approximately **2 s at 500 Hz**, **history length** of **64 frames**, **number of controlled modes** equal to **209**, and **planning horizon** $H = 4$ [2606.10771]. The GHOST paper states that performance is especially sensitive to the **history length** and **planning horizon**, which must align with the system delay and dominant disturbance periods [2401.00242]. This is consistent with the role of PO4AO as a predictive controller: too short a history or too short a simulated horizon would undercut its ability to learn temporal structure.

Real-time feasibility is a recurring engineering issue. The 2022 study reports **inference time** of roughly **0.3 ms** per step for an 8-m simulation and **0.35–0.7 ms** for the ELT-scale case on GPU hardware [2205.07554]. The GHOST implementation introduced approximately **700 microseconds** of additional latency beyond hardware, pipeline, and Python interface latency [2401.00242]. In the on-sky Papyrus deployment, the Python implementation introduced approximately **720 microseconds mean latency** with **$\sigma \approx 480$ microseconds** and up to **5% frame drops**, while the paper abstract summarizes this as approximately **$750\,\mu\text{s}$ of additional latency**, plus control jitter and occasional frame drops [2606.10771]. The authors nonetheless report performance gains despite these prototype-level software constraints, and they identify lower-level implementations such as **C/C++**, **CUDA**, or **TensorRT** as paths toward lower latency [2401.00242] [2606.10771].

## 4. Simulation and laboratory validation

The first detailed evaluations of PO4AO were conducted in **numerical simulation** and **laboratory experiments** rather than on sky. The 2022 study used **COMPASS** simulations for a VLT-scale system with **1364 DM actuators** and a **$41\times41$ PWFS**, and an ELT-scale system with **over 10,000 actuators** and a **$121\times121$ PWFS** [2205.07554]. In those simulations, PO4AO was reported to improve **post-coronagraphic contrast** by factors of **4–7** over the integrator for a **0-mag** star with a non-modulated PWFS, by **3–9** under noisier **9 mag** conditions, and by **20–90** in an ideal WFS case [2205.07554]. The same work states that PO4AO could be trained on timescales of **5–10 seconds** and that inference remained below **1 ms**, which was presented as sufficient for real-time XAO control [2205.07554].

Laboratory demonstrations followed on **MagAO-X** and later **GHOST**. The 2022 paper reports that on the MagAO-X laboratory setup, PO4AO reduced **speckle variance** in science images by factors of **3–7** at **$2.4$–$6\,\lambda/D$** relative to the integrator [2205.07554]. The 2023/2024 GHOST study emphasized different strengths: transparent adaptation to increasing, unknown control-loop delays of **0**, **1**, or **2** full frames, strong low-flux performance down to **S/N $\sim 2$**, and robustness under **dynamic misregistration** with the DM shifted by up to **40% of actuator spacing** [2401.00242]. In that work, PO4AO outperformed the **optimally tuned integrator by a factor $\sim 3$ in wavefront residuals** in the time-delay experiments and remained stable without recalibration after system changes that destabilized a classical integrator unless recalibrated [2401.00242].

These studies also provide a mechanistic interpretation of PO4AO’s gains. The method is described as **predictive**, **self-calibrating**, and capable of handling **non-linear wavefront sensing**, including the **Pyramid WFS optical gain effect** [2205.07554] [2401.00242]. The GHOST experiments showed that longer history improves control performance and enables damping of periodic disturbances, including a **16 Hz vibration**, which supports the claim that PO4AO uses temporal windows not just for noise averaging but for learned predictive rejection of structured disturbances [2401.00242].

## 5. On-sky demonstration on Papyrus

The first on-sky validation of PO4AO was reported in 2026 and constitutes the first **on-sky demonstration of a reinforcement learning controller for adaptive optics** [2606.10771]. The controller was deployed on the **Papyrus adaptive optics system** installed at the **Coudé focus** of the **1.52 m telescope (T152)** at the **OHP**. Papyrus uses a **pyramid WFS**, a **241-actuator DM**, an IR science path, and a loop frequency of **500 Hz** [2606.10771].

The on-sky experiments compared PO4AO against a **standard integrator controller** over several nights and across different guide-star magnitudes, flux levels, seeing conditions, and the presence of telescope vibrations [2606.10771]. The study reports that PO4AO **consistently outperformed the standard integrator in all tested configurations** [2606.10771]. Strehl ratios extracted in the paper are summarized below.

| Target | PO4AO Strehl (%) | Integrator Strehl (%) |
|---|---:|---:|
| Vega | 12.1 | 10.5 |
| Vega (2nd night) | 26.5 | 18.2 |
| HD 177809 | 22.45 | 12.86 |
| Cygni | 18.9 | 9.9 |
| Alpha Cas | 13.3 | 8.0 |

The paper also reports **relative peak intensity** ratios of **0.77**, **0.74**, **0.68**, **0.60**, and **0.62** for the integrator relative to PO4AO in those configurations, again favoring PO4AO [2606.10771]. Under a night affected by a malfunctioning tracking motor, PO4AO reduced the **tilt-mode vibration PSD peak by a factor 4.5 at 80 Hz**, whereas the integrator could not [2606.10771]. Residual variance for most modes was lower by factors of **1.2–2.3**, with the largest gains on **low-order modes** [2606.10771]. Under poor seeing, reported as **2.2″ to 3.8″** in general and up to **4″** on the third night, PO4AO still delivered noticeably higher Strehl and peak intensity [2606.10771].

A notable operational conclusion from the on-sky study is that once the controller had been tuned for Papyrus, it used the same hyperparameters for all tested targets and conditions [2606.10771]. This is the most concrete basis in the literature for the description of PO4AO as a **turnkey** AO controller.

## 6. Distinctive properties, comparisons, and limitations

The PO4AO papers consistently position the method against the **classical integrator**, and, more broadly, against **Kalman/LQG** and **model-free RL** approaches [2401.00242]. The stated advantages include **predictive capability**, **self-calibration**, robustness to **misregistration**, tolerance of **photon noise**, and operation under **nonlinear** sensing conditions [2205.07554] [2401.00242] [2606.10771]. In the comparison table given in the GHOST study, PO4AO is characterized as **nonlinear, data-driven, robust**, able to **self-calibrate**, able to use **arbitrary history via neural networks**, and fast in online adaptation relative to classical alternatives [2401.00242].

Several features recur as technically distinctive.

First, PO4AO learns both **control** and **reconstruction from data**, rather than relying on a fixed interaction matrix and manually tuned gains [2401.00242]. Second, it performs **continuous adaptation** through online learning, allowing it to track changes in wind, optical gain, misregistration, and noise level [2205.07554] [2401.00242]. Third, it uses temporal history explicitly, which enables learned compensation for latency and periodic disturbances [2401.00242] [2606.10771]. Fourth, it can act as a **drop-in replacement** within existing AO data flow, since the real-time controller still provides WFS preprocessing and command transport while PO4AO supplies the control law [2205.07554] [2401.00242].

At the same time, the papers do not present PO4AO as free of engineering constraints. All three major studies emphasize latency sources introduced by the machine-learning implementation, especially **Python overhead**, **CPU-GPU synchronization**, and contention between inference and training threads [2401.00242] [2606.10771]. The on-sky deployment succeeded despite additional latency and frame drops, but the paper explicitly states that performance was achieved in spite of a **non-optimized Python implementation** [2606.10771]. This indicates that the observed performance should not be conflated with an upper bound on the method; rather, the authors argue that optimization could further improve results.

A common misconception is that reinforcement learning controllers for AO had remained purely simulated or laboratory-based. That characterization became outdated with the Papyrus result, which the 2026 paper presents as the first on-sky demonstration of such a controller [2606.10771]. Another misconception is that PO4AO requires per-target manual retuning; the on-sky study reports the opposite for the tested Papyrus configurations [2606.10771]. A more defensible caution is that turnkey operation was demonstrated **after** tuning for a specific system, not as a guarantee of zero commissioning effort across arbitrary AO architectures.

## 7. Prospects and scope

The current literature presents PO4AO as a controller designed for **single-conjugate AO** and **XAO** use cases, with a clear emphasis on **high-contrast exoplanet imaging** [2205.07554] [2606.10771]. Its demonstrated scope ranges from laboratory systems such as **MagAO-X** and **GHOST** to on-sky operation on **Papyrus** [2205.07554] [2401.00242] [2606.10771]. The 2026 paper proposes broader deployment to more challenging and higher-order systems, including **MagAO-X**, **SCExAO**, and upgraded **VLT/SPHERE/SAXO+**, and identifies extension to more nonlinear wavefront sensing scenarios, such as **Zernike** and **non-modulated pyramid WFS**, as future directions [2606.10771].

The method is especially relevant to regimes where existing AO controllers encounter known failure modes: **delay-dominated control**, **non-stationary calibration**, **vibration contamination**, **low-flux operation**, and **dynamic misregistration** [2205.07554] [2401.00242] [2606.10771]. The 2022 simulations further suggest that the approach scales to **ELT-size systems**, with sub-millisecond inference reported for configurations exceeding **10,000 actuators** [2205.07554]. The 2026 paper adds that future work should examine retraining frequency, parallelization strategies, and real-time determinism for systems with **5,000–10,000+ actuators** [2606.10771].

Taken together, the published results define PO4AO as a model-based RL controller that learns a temporal-spatial control policy directly from AO telemetry, uses an internal dynamics model to optimize that policy over a short horizon, and has now been validated from simulation through laboratory operation to on-sky deployment [2205.07554] [2401.00242] [2606.10771]. The available evidence supports its characterization as a predictive, self-calibrating, and system-compatible AO control framework, while also indicating that software optimization remains a significant part of its maturation path.

Source: https://www.emergentmind.com/topics/po4ao