---
title: Squint in Wireless, Learning & Robotics
url: https://www.emergentmind.com/topics/squint
type: topic
---

# Squint in Wireless, Learning & Robotics

Squint denotes several technical concepts in current arXiv literature. In wireless communications it most often refers to **beam squint**, the frequency-dependent shift of a beam’s main lobe or focal point in wideband arrays. In online learning it names a **second-order, parameter-free algorithm** for prediction with expert advice and later extensions for changing environments and global second-order control. In robotics it names a **visual Soft Actor–Critic method** engineered for fast wall-clock training and zero-shot sim-to-real transfer [1705.04441, 2209.06826, 2602.21203].

## 1. Technical senses of the term

In the literature represented here, the term is used in three distinct ways.

| Domain | Meaning | Representative arXiv source |
|---|---|---|
| Wideband wireless/array processing | Beam squint: frequency-dependent beam or focus displacement | [1705.04441] |
| Online learning | Squint: second-order expert-advice algorithm and later variants | [2209.06826] |
| Sim-to-real robotics | Squint: visual SAC method with “resolution squinting” | [2602.21203] |

The first usage is by far the broadest. It spans switched-beam codebooks, hybrid beamforming, THz and sub-THz systems, near-field XL-MIMO, IRS design, integrated sensing and communications, and wideband OTFS. The second usage is specific to sequential decision-making with expert advice, where Squint is defined through a mixture over learning rates and a second-order potential. The third usage is a proper algorithm name in reinforcement learning, where “resolution squinting” denotes a deliberate render-then-downsample observation pipeline rather than any electromagnetic effect [1705.04441, 2603.03409, 2602.21203].

## 2. Beam squint as a wideband array phenomenon

In phased arrays, beam squint arises because the phase shifters are typically fixed for the carrier frequency and do not realize the exact time delays required at other frequencies. For a ULA with \(N\) antennas, spacing \(d\), incident angle \(\theta\), and frequency \(f\), the steering vector is
\[
a(f,\theta)=\bigl[1,e^{j2\pi f(d\sin\theta)/c},\dots,e^{j2\pi f(N-1)(d\sin\theta)/c}\bigr]^T.
\]
If the beamformer is fixed at \(f_c\) to focus at \(\theta_F\), the phase shifts are
\[
\beta_n=-2\pi f_c\,d\,(n-1)\sin\theta_F/c.
\]
Using the virtual-angle notation \(\psi=\sin\theta\) and \(\xi=f/f_c\), the maximum gain occurs when \(\xi\psi=\psi_F\), so the effective pointing direction is \(\psi=\psi_F/\xi\), and the squint in virtual angle is
\[
\Delta\psi_{\text{squint}}(\xi)=\psi_F-\psi_F/\xi=\psi_F(1-1/\xi).
\]
For small fractional offset, \(\Delta\psi_{\text{squint}}\approx (b/2)\psi_F\), with \(b=B/f_c\) [1705.04441].

An equivalent formulation in wideband OFDM uses the subcarrier frequencies
\[
f_m=f_c+\frac{B}{M}\Bigl(m-1-\frac{M-1}{2}\Bigr),
\]
and the equivalent spatial angle
\[
\varphi_m=\frac{f_m}{f_c}\phi.
\]
The dimensionless squint factor is
\[
\varepsilon=\frac{B}{2f_c},
\]
so \(\varphi_m\in[(1-\varepsilon)\phi,(1+\varepsilon)\phi]\). This makes explicit that wider bandwidths cause different subcarriers to “see” different steering directions even when the array geometry is fixed [2101.06845].

The same mechanism appears beyond ULAs. In UPAs, the normalized array gain becomes the product of two Dirichlet terms, one per spatial dimension, and the paper on THz communications summarizes the severity through a Beam Squint Ratio,
\[
\mathrm{BSR}\approx \frac{b}{8}\max\{N_{\rm r,h}\Delta_{\rm r,h},N_{\rm r,v}\Delta_{\rm r,v}\}.
\]
With half-wavelength spacing and fixed total \(N_r\), the BSR is minimized by a square UPA, and the paper states
\[
\mathrm{BSR}_{\rm upa}=\frac{b\sqrt{N_r}}{16},\qquad
\mathrm{BSR}_{\rm ula}=\frac{bN_r}{16}.
\]
In a separate wideband large-scale MIMO analysis, the closed-form beam-squint ratio for a ULA is
\[
\mathrm{BSR}\approx \frac{N\,b\,\Delta}{8},
\]
which again scales linearly with antenna count and fractional bandwidth [2303.12466, 2210.06890].

Near-field wideband systems generalize angular squint to **spatial** squint. A phase-shifter vector that focuses at \((r,\theta)\) for \(f_c\) only perfectly focuses the central subcarrier; the remaining subcarriers shift to different \((r+\Delta r_m,\theta+\Delta\theta_m)\). In near-field ISAC work this is described as a continuous spatial trajectory of beam foci across subcarriers, and in wideband XL-MIMO it is the basis for controllable beam squint with true-time-delay lines [2309.14012, 2412.01029].

Performance degradation follows directly from this frequency dependence. For a single-path OFDM channel, the wideband capacity with beam squint,
\[
C_{\mathrm{BS}}(\psi_F,\psi,b)=\frac{B}{N_f}\sum \log_2\!\left[1+\frac{P|g(\xi_n\psi-\psi_F)|^2}{B\sigma^2}\right],
\]
satisfies
\[
C_{\mathrm{BS}}(\psi_F,\psi,b)\le C_{\mathrm{NBS}}(\psi_F,\psi),
\]
with strict inequality except at broadside. In fixed-size codebooks with idealized beams, the average spectral efficiency for small \(\varepsilon\le 2/L\) satisfies
\[
\bar R\approx \log_2(1+\rho L)\Bigl[1-\Bigl(\frac14+\frac{L}{8}\Bigr)\varepsilon\Bigr],
\]
so the drop is linear in the squint factor [1705.04441, 2101.06845].

## 3. Compensation and suppression in communication systems

A first line of work treats beam squint as a codebook-design problem. In switched-beam systems, each beam \(w_i\) is assigned a coverage set
\[
\mathcal R_c(w_i,C_t)=\{\psi\in[-\psi_m,\psi_m]: C_{\mathrm{BS}}(\psi_{i,F},\psi,b)\ge C_t\},
\]
and the objective is to minimize codebook size subject to complete angular coverage. The resulting beam-alignment procedure starts at broadside, finds coverage edges by binary search, mirrors beams symmetrically, and continues until the target interval is covered. The paper reports that, for \(N=64\), \(f_c=73\) GHz, and \(B=2.5\) GHz, the squint-aware design yields up to \(17.8\%\) higher minimum capacity; at the same operating point, roughly \(M\approx 40\) beams are needed rather than \(M\approx 32\) if squint is ignored. It also identifies a supremum \(b_{\rm sup}\approx a/N\) beyond which no finite codebook can satisfy the capacity constraint [1705.04441].

A second line assumes the codebook size is fixed and optimizes beam shapes. One formulation samples the expanded angular interval seen under squint, introduces weights \(t(\varphi^k)\), and maximizes a weighted sum of \(\log_2(1+\rho r_k)\) subject to \(|a(\varphi^k)^H w|^2\ge r_k\). Because these constraints are non-convex, the paper applies the Concave–Convex Procedure, linearizes each quadratic form around the current iterate, and solves the resulting convex program in CVX. The reported outcome is that enlarged coverage alone only partially recovers edge-subcarrier rates, whereas the full optimization slows the spectral-efficiency degradation; at \(\varepsilon=0.05\), the proposed codebook recovers \(\sim 20\%\) more spectral efficiency than DFT, and the design guideline given is \(2/L\approx \varepsilon\) at worst-case bandwidth [2101.06845].

Hybrid beamforming generalizes the mitigation problem to array architecture. For THz UPAs, one proposed design builds the frequency-flat analog combiner from the dominant eigenvectors of the subcarrier-averaged sample covariance and then applies phase-only projection. Because the analog part is derived from all subcarriers rather than only the carrier, it is less sensitive to squint. The same study emphasizes that a square UPA is intrinsically more robust than a ULA of the same aperture. In switch-based HBF, the reported contrast is sharper: the expected array gain of PS-based beamforming decreases monotonically with BSR and approaches \(1/3\), whereas the switch-based approximation is
\[
\bar g_{\rm sw}\approx \frac{2}{3}\sqrt{\frac{\|w\|_1}{N}},
\]
which stays above \(1/3\) when many antennas are connected. Under \(f_c=300\) GHz, \(N=256\), \(N_{\rm RF}=4\), and bandwidth up to \(41.25\) GHz, the proposed SW-HBF is reported to achieve over \(38\%\) higher spectral efficiency than \(2\)-bit FC-PS-HBF and energy-efficiency gains exceeding \(60\%\) at \(20\) dB in the single-user case [2303.12466, 2210.06890].

Other mitigations use diversity, delay elements, or geometric reconfiguration. A constant-modulus beamformer designed by semidefinite relaxation can switch between direct beamforming and Alamouti STBC based on the rank structure of the relaxed solution; at \(f_c=28\) GHz and \(M=64\), the reported band-edge gain loss is reduced from \(\approx 31\) dB to \(\approx 5\) dB, with throughput gains of \(19\%\) at \(2\) GHz bandwidth and \(24.6\%\) at \(3\) GHz bandwidth. Delay-adjustable metasurfaces for THz IRS communications choose per-element phase shifts and time delays to cancel the affine-in-frequency phase terms; in the representative \(f_c=200\) GHz, \(B=6\) GHz, \(R=64\) setting, the beam-gain ripple is reported as less than \(2\%\) in the far field and restored to within \(5\%\) of the peak in the near field. Wideband near-field suppression via movable antennas formulates a max–min analog-gain problem over antenna positions and solves it with an SGDA-based block-coordinate procedure, producing an almost perfectly flat gain curve across \(30\)–\(50\) GHz. A related THz design jointly optimizes analog beamforming and 3D rotation; its reported minimum-gain improvement is \(\approx 15\) dB versus no rotation [1808.10117, 2208.12385, 2407.19511, 2503.08134].

Beam squint can also couple with other wideband impairments. In massive MIMO-OTFS, beam squint and Doppler squint form a doubly-squint effect. The cited work derives a peak-index-based channel estimator and a hybrid precoder with TTD, phase-shifter, and OTFS-domain compensation, and reports that the proposed design outperforms Doppler-only, delay–phase, and conventional PDMA precoding in achievable delay–Doppler-grid rate [2504.08569].

## 4. Beam squint as a sensing and localization resource

A recurrent theme in recent ISAC work is that beam squint need not only be mitigated; it can be engineered as a frequency-domain scanner. In a wideband massive-MIMO OFDM system with TTD lines, one design chooses the phase-shifter setting \(\phi\) and delays \(t_m\) so that the beam points to \(\theta_0\) at \(f=0\) and to \(\theta_c\) at \(f=F\). The resulting main-lobe direction satisfies
\[
\sin\theta(f)=\frac{\sin\theta_0+(f/F)\bigl((1+F/f_n)\sin\theta_c-\sin\theta_0\bigr)}{1+f/f_n},
\]
so the subcarriers sweep monotonically across the desired angular interval. If inter-element spacing is enlarged beyond \(\lambda_n/2\), beam split creates additional lobes and expands the sensing range. The claimed operational consequence is that \(Q\approx N\) frequency-domain beams can be transmitted within a single OFDM symbol, reducing over-the-air training time by roughly a factor of \(Q\); with one extra intersection repeat for beam-split disambiguation, only two OFDM symbols are needed [2207.08737].

Near-field ISAC uses an analogous idea in joint angle–range space. For a phase-only near-field beamformer designed at \((r_0,\theta_0)\) and reference frequency \(f_0\), the squinted focus on subcarrier \(f_m\) follows
\[
\sin\theta_m=\frac{f_m}{f_0}\sin\theta_0,\qquad
r_m=\frac{f_m r_0\cos^2\theta_m}{f_0\cos^2\theta_0}.
\]
With TTDs, the start and end points of the trajectory can be fixed at chosen anchors \((r_0,\theta_0)\) and \((r_c,\theta_c)\), allowing the system to “draw” a trajectory of beam foci across the near-field volume. Localization then reduces to identifying the peak-power subcarrier and, in higher-accuracy variants, combining multiple sweeps and phase differences. The reported CBS-Low method reduces beam sweeps by \(>99.9\%\), while CBS-High with \(P=5\) sweeps yields angle RMSE \(\sim 0.02^\circ\) and range RMSE \(\sim 0.04\) m at \(10\) dB SNR [2309.14012].

Wideband XL-MIMO localization extends this idea further by combining controllable beam squint and deep learning. The cited formulation derives CRBs for joint angle–range estimation under spatial non-stationarity, then proposes a three-stage CBS-based beam-training procedure: coarse angle, angular refinement by subcarrier grouping, and iterative range refinement. A ConvNeXt model then consumes the measurements and coarse estimates and regresses \((r,\theta)\) directly. The reported performance is centimeter-level accuracy, with \(\mathrm{RMSE}_r\approx 4\) cm and \(\mathrm{RMSE}_\theta\approx 10^{-3}\) rad at \(20\) dB [2412.01029].

These results correct a common oversimplification in beam-squint discussions. The phenomenon is indeed harmful to communication gain and rate under carrier-designed phase-only beamforming, but the same frequency dependence becomes informative when subcarriers are deliberately assigned distinct directions or focal points. This is the central methodological bridge between the mitigation literature and the sensing/localization literature [2207.08737, 2309.14012, 2412.01029].

## 5. Squint in online learning with expert advice

In online learning, Squint is a second-order algorithm for the expert problem. At round \(t\), the learner chooses \(p_t\in\Delta_N\), observes losses \(\ell_t\in[0,1]^N\), and incurs \(\langle p_t,\ell_t\rangle\). The instantaneous regret to expert \(i\) is
\[
r_{t,i}=\langle p_t,\ell_t\rangle-\ell_{t,i},
\]
with cumulative regret \(R_{t,i}=\sum_{s=1}^t r_{s,i}\) and second-order term \(V_{t,i}=\sum_{s=1}^t r_{s,i}^2\). The defining Squint potential is
\[
\Phi(R,V)=\int_0^{1/2}\frac{e^{\eta R-\eta^2V}-1}{\eta}\,d\eta,
\]
and the original update is
\[
p_{t,i}\propto \partial_R\Phi(R_{t-1,i},V_{t-1,i})
=\int_0^{1/2} e^{\eta R_{t-1,i}-\eta^2V_{t-1,i}}\,d\eta.
\]
As summarized in later notes, this produces a simultaneous \(\varepsilon\)-quantile regret guarantee in terms of the variance of the \(\varepsilon\)-quantile expert [2603.03409].

The 2022 changing-environment extension, Squint-CE, begins from the observation that a conventional black-box meta-wrapper destroys Squint’s favorable second-order behavior: the induced overhead \(\sqrt{|I|\ln I_2}\) dominates the sublinear advantages coming from variance adaptation. Squint-CE therefore intertwines Squint’s surrogate reduction with a single layer of exponential-weights meta-combination over geometric intervals. For every contiguous interval \(I\), it guarantees
\[
R_I^{\mathcal K}\le 2\sqrt{2V_I^{\mathcal K}A_I^{\mathcal K}}+4A_I^{\mathcal K},
\]
where
\[
A_I^{\mathcal K}
=
\Bigl(
2\lceil\log_2(|I|+2)\rceil\,
(\ln(2T)+\ln\lceil\log_2\sqrt T\rceil-\ln\pi(\mathcal K))
\Bigr)\vee 1.
\]
In big-\(O\) form, the bound is
\[
O\!\Bigl(\sqrt{\ln|I|\,V_I^{\mathcal K}\,(\ln T+\ln\ln T-\ln\pi(\mathcal K))}
+\ln|I|\,(\ln T+\ln\ln T-\ln\pi(\mathcal K))\Bigr),
\]
so the changing-environment version preserves second-order dependence on interval variance up to logarithmic factors [2209.06826].

A 2026 note proposes a simple variant that replaces the expert-specific \(V_{t,i}\) by a single global \(V_t\). The algorithm still computes
\[
w_{t,i}=\partial_R\Phi(R_{t-1,i},V_{t-1}),
\]
but then updates \(V_t=V_{t-1}+v_t\), where \(v_t\in[0,1]\) is chosen as the root of
\[
f(v)=\sum_{i=1}^N q_{t,i}(v-r_{t,i}^2)=0,
\]
with
\[
q_{t,i}\propto \partial_R^2\Phi(R_{t,i},V_{t-1}+v).
\]
The resulting \(\varepsilon\)-quantile regret bound has the same form as the original except that it depends on the global \(V_T\) rather than \(V_{T,i_\varepsilon}\), and the note states that it resembles the guarantee obtained by Freund et al. for a variant of NormalHedge [2603.03409].

Within this literature, the main conceptual distinction is therefore between **per-expert** and **global** second-order control, and between **static** and **changing-environment** regret. The term “Squint” refers to the family of algorithms built around the same potential-based, learning-rate-mixture construction rather than to a single fixed update rule [2209.06826, 2603.03409].

## 6. Squint in fast visual reinforcement learning for robotics

In robotics, Squint is an off-policy, vision-based actor–critic algorithm built on Soft Actor–Critic and designed to minimize wall-clock training time in massively parallel GPU simulation while transferring zero-shot to a real \(5\) DoF SO-101 robot arm. The high-level loop uses \(N=1024\) parallel ManiSkill3 environments at \(10\) Hz, renders wrist-camera RGB images at \(128\times 128\), downsamples them to \(16\times 16\), appends proprioception, and stores transitions in a GPU-resident replay buffer of size \(1\) M. After each environment step it performs \(U=256\) gradient updates of a shared two-layer CNN encoder, two C51-style distributional critics and their EMA targets, a stochastic Gaussian policy, and an entropy temperature \(\alpha\) [2602.21203].

The defining design choice is **resolution squinting**. The observation is not rendered directly at low resolution; instead the simulator renders at \(128\times128\) and applies area-interpolation downsampling to \(16\times16\),
\[
\hat o_t=D_{16\times 16}[R_h(\mathrm{env}_t)].
\]
The paper attributes two effects to this pipeline: lower compute for the two-layer CNN encoder and natural anti-aliasing that preserves object shape under heavy domain randomization. This component is coupled with LayerNorm after every linear layer in actor and critic heads, a tuned update-to-data ratio of roughly \(0.25\), and a systems stack using `torch.compile`, CUDA Graphs, mixed-precision \(bfloat16\) convolutions, and an entirely on-GPU replay buffer. The implementation is reported to achieve more than a \(5\times\) end-to-end speed-up over a naive off-policy visual agent [2602.21203].

The learning objectives retain standard SAC structure but add a distributional critic loss. The critic minimizes a soft Bellman residual, the actor minimizes
\[
L_\pi(\phi)=
\mathbb E[\alpha\log\pi_\phi(a|s)-\tfrac12(Q_{\theta_1}(s,a)+Q_{\theta_2}(s,a))],
\]
the temperature is adapted with a target-entropy loss, and each critic also minimizes a C51 cross-entropy to a projected Bellman target distribution. The encoder itself is deliberately small: two \(3\times3\) convolutional layers with channels \(16\to 32\to 64\), ReLU activations, and batch size \(512\) [2602.21203].

Empirically, Squint is evaluated on eight SO-101 manipulation tasks—Reach Cube, Reach Can, Lift Cube, Lift Can, Place Cube, Place Can, Stack Cube, Stack Can—with heavy visual and physical domain randomization. Policies are trained for \(15\) minutes on a single NVIDIA RTX 3090 GPU. The reported simulation mean success rate over all eight tasks after \(15\) minutes is \(96.1\%\pm 3.2\%\), and most tasks converge in under \(6\) minutes. In the real world, the zero-shot success rate across \(80\) trials is \(91.3\%\), which the paper describes as a \(25\%\) absolute improvement over state-to-visual DAgger at \(66.3\%\) when accounting for the time to train its state-based expert. A visual robustness ablation reports that removing color jitter reduces success from \(91.3\%\) to \(72.5\%\) [2602.21203].

In this usage, “Squint” is not related to beam steering or expert-advice regret. It is the proper name of a visual SAC system whose central innovations are parallel simulation, a distributional critic, anti-aliased low-resolution observations, tuned update-to-data ratio, and GPU-level implementation choices [2602.21203].

Source: https://www.emergentmind.com/topics/squint