OTF-LAM: Latent Actions & Lens Aberrations
- The paper presents a framework that factorizes observed visual transitions into reusable motion primitives, achieving robust latent action learning with minimal degradation under ambiguity.
- OTF-LAM in optics leverages autocorrelation of the generalized pupil function and Zernike polynomial expansions to quantify lens aberrations via MTF and PTF metrics.
- Both approaches decompose complex transformations into interpretable factors, improving policy learning in robotics and enabling rapid non-Fourier imaging system analysis.
OTF-LAM refers to two fundamentally distinct domains: (1) Observed Transition Factorization–Latent Action Models in the context of representation learning for visual agents (Nam et al., 29 Jun 2026), and (2) an Optical Transfer Function–Lens Aberration Mechanism in physical optics (Bai et al., 2018). Both leverage the idea of decomposing or analyzing transitions or transformations for more robust modeling, but apply to separate technical problems.
1. Latent Actions from Factorized Transitions under Agent Ambiguity (OTF-LAM)
OTF-LAM is a framework that addresses learning robust action-like representations (latents) from observed transitions in visual domains, especially when agent influence is ambiguous due to distractors, camera motion, or background dynamics. Unlike traditional Latent Action Models, which compress entire observation transitions into a single latent, OTF-LAM first factorizes the observed transition into a set of elementary motion-effect primitives before learning action latents.
Key Stages and Methods
- Transition Preprocessing: Each observation transition is pre-processed via edge or gradient filtering (e.g., using Sobel operators) to produce motion-centered differences, such as first-order velocity or second-order acceleration across frames. These motion signals retain both agent and non-agent induced effects.
- Sparse Patchwise Factorization: Motion representations are split into non-overlapping patches, each embedded through a small MLP and quantized using a VQ codebook into discrete motion primitives. Each primitive is described by its code, a spatial occupancy map indicating which patches activate the code, and an activation strength. This sparse, reusable vocabulary of “observed-transition factors” serves as a compositional basis for subsequent processing.
- Aggregation to Latent Action: For each active primitive, a state-aware token is synthesized and modulated by a gating weight. These are aggregated to form a state-controlled factor-level vector, which is mapped to the final action-like latent embedding for downstream use.
- Inverse-Forward Dynamics Training: The model is trained via an inverse–forward dynamics paradigm. An inverse head predicts the agent action given both current and next state encodings, while a forward model predicts the next state encoding from the current state and latent action, using cross-entropy and mean-squared-error losses, respectively. The factorization mapping is learned on frozen, pre-trained transition vocabularies.
- Decoder-Free Variant (OTF-LAM-Dino): This variant uses a frozen DINOv2 image encoder and performs prediction in the latent feature space rather than reconstructing pixels, further removing pixel-level nuisance variation from the learning target.
Core Equations
- Patch quantization:
- Sparse decomposition:
- Aggregation and latent action formation:
- Inverse and forward losses:
Empirical Findings
- Zero-shot Transfer: Factorized primitives transfer across changes in observed carriers and agent morphologies with minimal performance loss (<25% degradation), whereas monolithic VQ-VAE approaches see substantially greater drops (60–70%).
- Policy Learning: Downstream behavioral cloning policies trained on the latent actions derived from factorized transitions match or outperform baselines, despite relying on a reusable, unsupervised codebook.
- Decoder-Free Robustness: Predictions in DINO feature space improve mean return by 10–30% and demonstrate greater insensitivity to pixel-level nuisance.
These results indicate that sparse factorization of observed transitions yields not only action representations with greater structure and transferability, but also significant robustness to visual clutter and ambiguity (Nam et al., 29 Jun 2026).
2. Optical Transfer Function–Lens Aberration Mechanism (OTF-LAM) in Physical Optics
In the context of optical system analysis, OTF-LAM refers to a domain-specific formulation of the Optical Transfer Function focusing on lens aberrations and their direct spatial-domain effects (Bai et al., 2018). Here, the OTF characterizes how an optical system transfers spatial frequency content from object to image, incorporating both amplitude attenuation (MTF) and phase distortion (PTF).
Fundamental Definitions and Derivations
- Cosine Fringe Test Pattern: An object-plane cosine fringe with spatial frequency is used as a basis to analyze response:
- Pupil Mapping and Aberration: The exit pupil of the lens contributes a phase factor via the wavefront error expanded in Zernike polynomials. The generalized pupil function is
- Interference and Modulation: The OTF is shown to be the autocorrelation of the generalized pupil function over the aperture:
which is decomposed into the Modulation Transfer Function (magnitude) and Phase Transfer Function (complex angle), tracing spatial-frequency dependent contrast and phase fidelity.
- Interpretation: Aberrations appear directly as phase variations in the pupil autocorrelation, and can be traced to loss of MTF (sharpness/contrast) and non-trivial PTF (spatial misalignment or asymmetric response).
Practical Significance
The OTF–LAM framework is advantageous for rapid, non-Fourier numerical computation of imaging system behavior under arbitrary aberrations, supporting tolerancing, lens design, and quality evaluation. By parameterizing the aberration phase via Zernike modes, one can predict degradation across the spectrum of spatial frequencies, enabling informed engineering decisions (Bai et al., 2018).
3. Comparative Table: OTF-LAM in Learning vs. Optics
| Framework | Domain | Core Mechanism |
|---|---|---|
| OTF-LAM (Nam et al., 29 Jun 2026) | Representation learning (AI/robotics) | Sparse factorization of observed state transitions for robust, transferable latent action modeling |
| OTF-LAM (Bai et al., 2018) | Physical optics | Autocorrelation of generalized pupil function for analyzing lens aberrations via OTF/MTF/PTF |
4. Significance and Impact
OTF-LAM in latent action learning provides a principled, unsupervised approach for extracting reusable transition primitives under ambiguous observational regimes—enabling robust policy learning with minimal supervision and strong transfer properties. In optical engineering, the OTF-LAM formalism supplies a fundamental, spatial-domain method to connect aberration physics to imaging quality rapidly and analytically, which is especially useful for complex or multi-element optical systems.
5. Related Work and Future Directions
In representation learning, OTF-LAM connects to object-centric, compositional, and disentangled latent variable models, improving upon prior methods by specifically targeting transition factorization before abstraction into agent-related latents. The introduction of the decoder-free DINO feature-space variant opens further avenues for research on invariant prediction targets and cross-domain transfer.
Within optics, alternative and non-Fourier OTF approaches continue to be explored for the analysis of complex, engineered apertures and adaptive optics—where real-time evaluation of system performance under perturbations and design tradeoffs is critical.
A plausible implication is that, despite arising in entirely distinct technical contexts, both OTF-LAM instantiations reflect a convergent emphasis on the decomposition of transformations into interpretable, reusable primitives or factors for enhanced analysis, learning, and system optimization.