- The paper demonstrates the on-sky deployment of PO4AO, a reinforcement learning controller that outperforms traditional integrator controls under varied atmospheric conditions.
- The paper reports higher Strehl ratios (up to 26.5%) and reduced modal variance, confirming superior performance in vibration rejection and noise robustness.
- The paper highlights PO4AO's predictive, history-dependent control law which adapts in real time, paving the way for scalable RL applications in ELT-class adaptive optics.
On-Sky Validation of Reinforcement Learning for Adaptive Optics Control: The PO4AO Controller on PAPYRUS
Introduction
The paper presents an empirical, on-sky validation of Policy Optimization for Adaptive Optics (PO4AO), a reinforcement learning (RL) controller, demonstrating its deployment and operational efficacy on the PAPYRUS AO platform using the 1.52m telescope at Observatory de Haute-Provence (OHP). The work rigorously evaluates PO4AO against classical integrator control under a spectrum of turbulence, vibration, and guide-star brightness conditions, reporting performance metrics including Strehl ratios, modal residuals, power spectral densities, and controller latency. The narrative places this effort in context with recent trends toward machine learning-based AO—especially RL and supervised neural strategies for model identification, nonlinear compensation, and predictive control—but distinguishes PO4AO by its full RL closure and turnkey robustness on real system telemetry (2606.10771).
PAPYRUS is a modular pyramid-wavefront sensing AO instrument designed to facilitate rapid prototyping and cross-validation of advanced control schemes. The bench comprises dual-wavelength calibration, common path, visible WFS, and NIR paths, with a high-resolution ALPAO DM and OCAM2K camera. The baseline controller is a leaky integrator operating in a reduced Karhunen-Loève modal space with calibration via actuator pokes Figure 1.

Figure 1: Schematic overview of the PAPYRUS bench architecture, showing the calibration, sensing, and science sub-paths for AO experimentation.
The Python-based interface enabled direct switching between PO4AO and integrator control for side-by-side telemetry acquisition, albeit inducing additional latency (720 μs mean with significant jitter due to GPU contention).
PO4AO RL Methodology
PO4AO integrates two convolutional neural networks: the dynamics model p^ω predicts next-step WFS measurements from DM command history, while the policy πθ maps current and past observations and actions to new DM increments. Training is conducted online via episodic rollouts on the dynamics model, with reward defined as negative squared residual amplitude post-reconstruction. The approach leverages dataset diversity induced by initial noise injection during warm-up, facilitating robust exploration and learning of both “good” and “bad” control regimes. After warm-up tuning, PO4AO operated with fixed hyperparameters throughout all on-sky tests. History depth and episode durations were empirically tuned for a balance between vibration discrimination and real-time stability.
Experimental Results and Comparative Analysis
Vibration Rejection: First Night
Under severe telescope tracking-induced vibration, PO4AO consistently provided higher Strehl ratios (12.14% vs. 10.46%) and superior PSF morphology relative to the integrator Figure 2. Modal variance analysis revealed pronounced reductions at low-order modes—most notably in “tilt”—with PO4AO maintaining a variance gain of 2.3 and suppressing PSD peaks by factors of 4.5–8 at vibration harmonics Figure 3.

Figure 2: Comparison of science camera PSFs under first-night vibration; PO4AO delivers sharper core and higher Strehl.


Figure 3: First night modal variance and tilt power spectral density; PO4AO demonstrates marked vibration cancellation and predictive adaptation.
Turnkey Robustness and Noise Immunity: Second Night
With vibration eliminated, PO4AO operated across a broad guide-star magnitude range (Vega to Cygni, V=0.09–6.66), maintaining superior Strehl ratios (up to 26.5% vs. 18.2% on Vega) and improved PSFs for faint targets Figure 4. Modal variance gains up to 3.5 were observed at low-order modes, attesting to noise robustness and gain optimization beyond fixed integrator settings Figure 5.

Figure 4: Second night PSF comparisons across target magnitudes; PO4AO outperforms integrator on sharpness and Strehl.



Figure 5: Second night telemetry, showing modal variance reduction for PO4AO across diverse targets.
Comparison to Optimized Low-Latency Integrator: Third Night
PO4AO demonstrated performance gains relative to a manually optimized, low-latency integrator, achieving 13.3% Strehl versus 8.0% under challenging conditions (Figure 6–8). Although telemetry modal variance improvement was less pronounced, PO4AO PSFs retained a sharper profile, suggesting qualitative advantages in spatial control law.


Figure 6: Third night PSFs under bad seeing; PO4AO delivers consistently higher Strehl.

Figure 7: Telemetry modal variance; PO4AO retains advantage in actuator control residuals.
Linear Analysis: Control Law Structure and Predictive Nature
Detailed linear approximations of PO4AO and integrator command matrices elucidate the structural distinctions. PO4AO control matrices exhibit significant off-diagonal and historical terms, indicating cross-actuator and time-lag integration, in contrast to the integrator’s strictly diagonal and instantaneous structure (Figure 8–10). Analysis of temporal contribution weights and eigenvalue spectra confirms PO4AO's history-dependent, spatially heterogeneous gain optimization (Figures 11–12).

Figure 8: Linear control matrices; PO4AO incorporates cross-actuator and time-lagged measurements.

Figure 9: Actuator contribution maps for present and lagged measurements; PO4AO demonstrates wind-oriented prediction.

Figure 10: Variance-weighted temporal history contributions; PO4AO exploits extended sequence memory for control.

Figure 11: Control matrix eigenvalue spectra; PO4AO implements mode-specific gain adaptation.
Implementation and Scalability Considerations
Despite additional latency from Python and GPU contention, PO4AO maintained loop stability and performance at 500 Hz. Scaling to ELT-class systems (up to 128×128 DMs) will demand further optimization of convolutional and modal filter layers and hardware acceleration, as matrix operations dominate training thread runtime.
Practical and Theoretical Implications
The validation of PO4AO as a turnkey controller represents a major step toward general-purpose RL for AO. The approach eliminates expert gain tuning, adapts online to changing atmospheric and instrumental conditions, and demonstrates tolerance to implementation-induced jitter and frame drops. Future extensions are likely to target high-contrast imaging pipelines, nonlinear sensing (e.g., non-modulated pyramid and Zernike WFS), and further integration with data-driven pre-/post-processing stages (including deep NN reconstructors, e.g. [landman2025making]). Additionally, the demonstrated predictive capacity and automatic gain selection support theoretical advances in adaptive spatiotemporal modeling for AO, particularly as system orders and dynamic range increase with ELT deployments.
Conclusion
The paper establishes PO4AO as a robust, high-performance RL controller for single-conjugate AO systems, validated on-sky across diverse conditions. The performance gains over legacy integrators—both in vibration rejection and noise robustness—are realized via history-dependent, spatially adaptive control laws. The results indicate RL control strategies are mature for operational deployment in astronomical AO, contingent on hardware optimization and further integration with advanced sensing paradigms. PO4AO will serve as a model for future RL-driven AO in extremely large telescopes and high-contrast exoplanet imaging (2606.10771).