Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ping-Pong Sampling in mmWave MIMO

Updated 20 May 2026
  • Ping-Pong sampling is a two-sided, alternating pilot protocol that autonomously refines beam alignment in mmWave MIMO channels without needing explicit feedback.
  • The approach employs LSTM-based recurrent neural networks and DNNs to iteratively update pilot vectors and optimize beamformers and RIS phases for enhanced SNR.
  • Empirical evaluations demonstrate that with only 8 pilot exchanges, the method reduces overhead and outperforms traditional channel estimation techniques, even under low SNR conditions.

Ping-Pong sampling refers to a two-sided, alternating pilot signaling and active sensing protocol for beam alignment in millimeter-wave (mmWave) MIMO systems, where the transmitter (Tx) and receiver (Rx) iteratively and autonomously adapt their measurement vectors over multiple rounds without the need for explicit feedback. This class of strategies is particularly suited for scenarios in which both ends of a link are equipped with large antenna arrays but only a limited number of RF chains, and the channel is modeled as high-dimensional and structured by multipath sparsity. The ping-pong protocol facilitates rapid, robust, and interpretable joint beam alignment, and can be further generalized to incorporate reconfigurable intelligent surfaces (RIS) for joint beam and reflection coefficient optimization (Jiang et al., 2023).

1. System Model and Problem Formulation

The basic system model considers a point-to-point mmWave MIMO channel with a transmitter (“agent A”, Tx) of NtN_t antennas and one RF chain, and a receiver (“agent B”, Rx) of NrN_r antennas and one RF chain, both configured as ULAs (Uniform Linear Arrays) with half-wavelength spacing. The narrowband, block-fading channel matrix from Rx to Tx is GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}, remaining constant over the full sensing and data transmission interval. The channel is represented as a sparse sum of a few multipath components: G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H, where αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1), at(),ar()\mathbf{a}_t(\cdot), \mathbf{a}_r(\cdot) are standard array response vectors, and LL is the number of paths. The objective is to identify transmit and receive beamformers (f,w)(\mathbf{f}^*, \mathbf{w}^*) that maximize the receive SNR: maxf=w=1wHGHf2.\max_{\|\mathbf f\|=\|\mathbf w\|=1} |\mathbf{w}^H\mathbf{G}^H\mathbf{f}|^2. With perfect CSI, the optimal solution is given by the principal singular vectors of G\mathbf{G}.

An RIS-assisted extension models the effective MIMO channel as

NrN_r0

where NrN_r1 and NrN_r2 represent the Tx–RIS and RIS–Rx sub-channels and NrN_r3 is the vector of unit-modulus RIS phase shifts.

2. Ping-Pong Pilot Transmission Protocol

The core of ping-pong sampling is an NrN_r4-round pilot exchange protocol. In each round NrN_r5, two steps occur:

  • Ping (Tx → Rx): The transmitter sends a scalar pilot NrN_r6 using pilot beamformer NrN_r7. The receiver observes:

NrN_r8

applying its current sensing vector NrN_r9.

  • Pong (Rx → Tx): Based on all observations GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}0, Rx selects its transmit sensing vector GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}1, transmits back a pilot, and Tx observes:

GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}2

using its current receive vector GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}3.

Each agent thus accumulates a personal sequence of pilot measurements and autonomously adapts its sensing protocol for future rounds. Notably, there is no explicit feedback channel—adaptation is based solely on local measurements. The formal protocol is captured by update rules (GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}4, GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}5, GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}6, GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}7), where each side generates its next sensing vector based on history.

3. Deep Active Sensing Framework

To operationalize the adaptive policy mappings in ping-pong sampling, each side employs a dedicated LSTM-based recurrent neural network. At each round:

  • The new scalar measurement (real and imaginary parts) updates the LSTM’s hidden and cell states via standard gate mechanisms.
  • These states parameterize small fully connected DNN “heads” which output the transmit and receive pilot vectors for the next round, enforced to have unit GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}8 norm.

After GCNt×Nr\mathbf{G} \in \mathbb{C}^{N_t \times N_r}9 rounds, the final LSTM cell states summarize the entire local history of observations. Two further DNNs (G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,0, G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,1) translate these cell states into final data phase beamformers (G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,2, G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,3). If an RIS is present, a third DNN (G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,4) produces the final RIS phase vector (G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,5), applying element-wise normalization to enforce the unit-modulus constraint.

All learnable parameters (LSTM, DNN heads) are trained end-to-end to maximize the expectation of the beamforming gain over random instantiations of G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,6 and noisy measurements: G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,7 Training is performed through stochastic gradient methods (SGD/Adam) over large Monte Carlo batches.

4. Beamformer Optimization and RIS Control

The ping-pong protocol can be directly compared to closed-form solutions available under perfect CSI. For the pure MIMO case, the benchmark is the SVD-based solution (G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,8). With RIS present, alternating optimization is employed: given RIS phases, apply SVD to G==1Lαat(ϕt)ar(ϕr)H,\mathbf{G} = \sum_{\ell=1}^L \alpha_\ell\,\mathbf{a}_t(\phi^{\rm t}_\ell)\mathbf{a}_r(\phi^{\rm r}_\ell)^H,9; with fixed beamformer pair, update each RIS phase to maximize alignment.

In the learned (active) scheme, all vectors are generated by the corresponding DNNs fed with the LSTM histories. In the RIS setting, the DNN for RIS control ingests the concatenated final LSTM cell states from both Tx and Rx.

5. Empirical Performance and Interpretability

The ping-pong sampling protocol, implemented via the deep active sensing architecture, demonstrates:

  • Pilot Overhead Reduction: With only αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1)0 ping-pong rounds (αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1)1 pilots), the approach surpasses classical compressive sensing plus SVD, and fixed/random DNN-based schemes using 20 pilots, at SNR = 0 dB.
  • SNR Robustness: Performance remains strong at SNR αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1)2 dB, outperforming standard channel estimation protocols measured at SNR = 0 dB.
  • Path Generalization: Networks trained on 3-path channels retain high performance with varying numbers of paths (1–5).
  • Interpretability: In early rounds, the learned sensing beams are broad, approximating uniform coverage; in later rounds, they narrow to concentrate on promising directions. Final learned data beams coincide with the principal singular vectors approximately 63% (top), 26% (second), etc., of the time for αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1)3 rounds.
  • RIS Adaptation: Joint design of RIS phases in the active sensing process delivers an additional 1–2 dB gain compared to static or random RIS phase codes.
  • Computational Complexity: At inference, each round involves an LSTM update and a few DNN forward passes, with total cost αCN(0,1)\alpha_\ell \sim \mathcal{CN}(0,1)4.

6. Applications and Extensions

Ping-pong sampling is immediately relevant for low-feedback, high-efficiency beam alignment in next-generation mmWave wireless systems, particularly where channel sparsity, rapid environmental dynamics, and limited RF chains render classical schemes inefficient. The approach’s adaptability and compatibility with RIS control broaden its utility to programmable and intelligent radio environments, supporting both directly connected and reflection-assisted links. A plausible implication is more robust operation under stringent pilot and SNR constraints, with direct impact on practical 5G/6G deployments (Jiang et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ping-Pong Sampling.