---
title: 'AirFC: Over-the-Air FC Layer Emulation'
url: https://www.emergentmind.com/topics/airfc
type: topic
---

# AirFC: Over-the-Air FC Layer Emulation

AirFC, short for **Air Fully Connected layers over the air**, denotes a computational paradigm in which a neural-network fully connected (FC) layer is realized by a RIS-assisted MIMO over-the-air computation system whose effective end-to-end wireless channel is configured to emulate the target FC weight matrix \(\mathbf W\) [2505.01170, 2508.01840]. In its canonical form, the digital layer
\[
\mathbf y=\mathbf W\mathbf x+\mathbf b
\]
is replaced by an analog transmission layer parameterized by a transmit precoder, one or more reconfigurable intelligent surfaces (RISs), and a receive combiner, so that wireless propagation itself functions as a programmable linear operator. Relative to standard over-the-air computation, which usually targets sums, means, or related aggregation primitives, AirFC is directed at FC-layer emulation and thereby at over-the-air neural inference [2508.01840].

## 1. Conceptual origin and problem setting

The AirFC formulation was introduced for RIS-assisted MIMO systems that emulate the functionality of a digital complex-valued FC layer by jointly configuring the wireless environment and the transceiver [2505.01170]. The central observation is that the superposition property of the wireless multiple-access channel can be used not only for analog aggregation but also for implementing matrix-vector multiplication when the propagation environment is sufficiently controllable.

The target layer is a complex FC map with input \(\mathbf x\in\mathbb C^{N\times 1}\), output \(\mathbf y\in\mathbb C^{N\times 1}\), weight matrix \(\mathbf W\in\mathbb C^{N\times N}\), and bias \(\mathbf b\in\mathbb C^{N\times 1}\). AirFC replaces this digital transformation with a RIS-aided transmission structure having \(N\) transmit antennas, \(N\) receive antennas, and either one RIS or multiple RISs. The key design task is to make the effective wireless operator approximate \(\mathbf W\) as closely as possible while accounting for noise, power limits, and unit-modulus RIS constraints [2505.01170].

This places AirFC at the intersection of analog over-the-air computation, programmable radio environments, and neural inference. A plausible implication is that the FC layer is no longer treated as a purely digital primitive but as a physically instantiated linear transform whose fidelity depends on channel structure, aperture, and hardware configurability.

## 2. RIS-assisted MIMO formulation

In the multi-RIS form, RIS \(i\) has \(M_i\) reflecting elements with
\[
\sum_{i=1}^L M_i=M,\qquad M_i=M/L.
\]
For RIS \(i\), the transmitter-to-RIS channel is \(\mathbf{\bar H}_i\), the RIS-to-receiver channel is \(\mathbf{\hat H}_i\), and the phase-shift matrix is
\[
\mathbf{\Theta}_i=\mathrm{diag}(e^{j\theta_{i,1}},\dots,e^{j\theta_{i,M/L}}).
\]
The direct Tx-Rx link is assumed blocked in the main model. The received signal is
\[
\mathbf y=\mathbf F_2\left(\sum_{i=1}^L \mathbf{\hat H}_i\mathbf{\Theta}_i\mathbf{\bar H}_i\mathbf F_1\mathbf x+\mathbf n\right)+\mathbf b,
\]
or equivalently
\[
\mathbf y=\mathbf F_2\left(\mathbf{\hat H}\mathbf{\Theta}\mathbf{\bar H}\mathbf F_1\mathbf x+\mathbf n\right)+\mathbf b,
\]
with \(\mathbf n\sim\mathcal{CN}(0,\sigma^2\mathbf I_N)\) [2505.01170].

The effective operator implemented by the wireless medium is therefore
\[
\mathbf F_2\left(\sum_{i=1}^L \mathbf{\hat H}_i\mathbf{\Theta}_i\mathbf{\bar H}_i\right)\mathbf F_1,
\]
and AirFC seeks to make this operator approximate \(\mathbf W\). The corresponding imitation-error minimization problem is
\[
\min_{\mathbf F_1,\mathbf F_2,\mathbf\Theta}
\left\|\mathbf F_2\mathbf{\hat H}\mathbf\Theta\mathbf{\bar H}\mathbf F_1-\mathbf W\right\|_F^2
+\mathbb E_{\mathbf n}\!\left\{\|\mathbf F_2\mathbf n\|^2\right\}
\]
subject to
\[
\|\mathbf F_1\|_F^2\le P_{\max},\qquad |\mathbf\Theta_{i,i}|=1,\; i=1,\ldots,M.
\]
The first term measures mismatch between the physical channel and the FC weights; the second penalizes noise propagation through the combiner [2508.01840].

## 3. Alternating optimization and closed-form block updates

Because \(\mathbf F_1\), \(\mathbf F_2\), and \(\mathbf\Theta\) are multiplicatively coupled and the RIS variables satisfy unit-modulus constraints, the AirFC design problem is non-convex. The proposed solution is a three-block alternating optimization procedure that sequentially updates the precoder, combiner, and RIS phases until convergence [2505.01170].

For fixed \(\mathbf F_2\) and \(\mathbf\Theta\), with
\[
\mathbf\Upsilon=\mathbf F_2\mathbf{\hat H}\mathbf\Theta\mathbf{\bar H},
\]
the precoder update is a QCQP whose semi-closed-form solution is
\[
\mathbf F_1^{\rm opt}=
\left(\mathbf\Upsilon^H\mathbf\Upsilon+\lambda\mathbf I_N\right)^{-1}\mathbf\Upsilon^H\mathbf W.
\]
If the unconstrained solution satisfies the power budget, then \(\lambda=0\); otherwise \(\lambda\) is obtained by bisection because \(\|\mathbf F_1^{\rm opt}\|_F^2\) decreases monotonically in \(\lambda\) [2505.01170].

For fixed \(\mathbf F_1\) and \(\mathbf\Theta\), defining
\[
\bar{\mathbf\Upsilon}=\mathbf{\hat H}\mathbf\Theta\mathbf{\bar H}\mathbf F_1,
\]
the combiner update has the closed form
\[
\mathbf F_2^{\rm opt}=
\left(\bar{\mathbf\Upsilon}\bar{\mathbf\Upsilon}^H+\sigma^2\mathbf I_N\right)^{-1}
\mathbf W\bar{\mathbf\Upsilon}^H.
\]
This is a regularized linear MMSE-type update [2508.01840].

For the RIS phase design, the diagonal entries of \(\mathbf\Theta\) are collected into
\[
\mathbf v=\left(\mathbf\Theta_{1,1},\ldots,\mathbf\Theta_{M,M}\right)^T,\qquad |v_i|=1.
\]
The RIS subproblem becomes
\[
\min_{\mathbf v}\;
\mathbf v^H\mathbf\Omega\mathbf v-2\Re\{\mathbf v^T\boldsymbol\varphi\}
+\operatorname{tr}(\mathbf W\mathbf W^H)
\quad\text{s.t. } |v_i|=1,
\]
with \(\mathbf\Omega\) and \(\boldsymbol\varphi\) determined by the current \(\mathbf F_1\), \(\mathbf F_2\), and channel matrices. A majorization-minimization surrogate yields the closed-form phase update
\[
\mathbf v^{\rm opt}
=
e^{j\arg\left((\lambda_{\max}\mathbf I_M-\mathbf\Omega)\mathbf v^r-\boldsymbol\varphi^*\right)},
\]
where \(\mathbf v^r\) is the current iterate and \(\lambda_{\max}\) is the largest eigenvalue of \(\mathbf\Omega\) [2505.01170].

The low-complexity characterization of AirFC derives from the fact that each block has either a closed-form solution or a semi-closed-form solution with only a scalar bisection search.

## 4. Trainable AirFC and CSI-dependent versus CSI-free training

A later extension studies AirFC as a trainable architecture in which \(\mathbf F_1\), \(\mathbf F_2\), and \(\mathbf\Theta\) are treated as learnable parameters rather than as a one-shot emulation of a fixed \(\mathbf W\) [2508.01840]. The loss is
\[
{\cal L}_{\rm loss}
=
-\sum_{i=1}^C p_i\log(\hat p_i)
-\lambda_p\min\{0,P_{\max}-P_{\rm Tx}\},
\]
where \(C\) is the number of classes, \(p_i\) the true label, \(\hat p_i\) the predicted softmax probability, and \(\lambda_p\) a penalty coefficient for transmit-power violations.

Two training strategies are considered. In **centralized training**, whichever terminal has CSI—the transmitter or the receiver—optimizes the AirFC parameters and then sends them to the other terminal over a control link. In **distributed training**, CSI acquisition is avoided by exploiting channel reciprocity and back-and-forth transmissions: the forward pass sends \(\mathbf x\), the receiver computes the loss gradient, and the backward pass returns gradient-related information over the reciprocal channel so that both ends update parameters locally [2508.01840].

The over-the-air gradient expressions include
\[
\frac{\partial {\cal L}_{\rm loss}}{\partial \mathbf F_2}
=
\frac{\partial {\cal L}_{\rm loss}}{\partial \mathbf y}
\left(\mathbf H_2\mathbf\Theta\mathbf H_1\mathbf F_1\mathbf x+\mathbf n\right)^T
\]
and
\[
\frac{\partial {\cal L}_{\rm loss}}{\partial \mathbf F_1}
=
\left(\mathbf F_2\mathbf H_2\mathbf\Theta\mathbf H_1\right)^T
\frac{\partial {\cal L}_{\rm loss}}{\partial \mathbf y}\mathbf x^T.
\]
The distributed updates are therefore noisy versions of the ideal gradients but eliminate explicit CSI estimation [2508.01840].

This extension preserves the original AirFC thesis: the wireless environment is treated as part of the model itself, not merely as a communications substrate.

## 5. Rank deficiency, multi-RIS design, and empirical behavior

A central limitation of single-RIS AirFC is rank deficiency in LoS-dominated channels. For a single RIS with LoS channels \(\mathbf H_1\) and \(\mathbf H_2\),
\[
\operatorname{rank}(\mathbf H_2\mathbf\Theta\mathbf H_1)
=
\min\{\operatorname{rank}(\mathbf H_1),\operatorname{rank}(\mathbf H_2)\}
=1\ll \operatorname{rank}(\mathbf W),
\]
so a single-RIS effective channel may be fundamentally unable to approximate a full-rank FC layer [2505.01170].

The multi-RIS extension addresses this by using
\[
\mathbf H=\sum_{i=1}^L \mathbf{\hat H}_i\mathbf\Theta_i\mathbf{\bar H}_i
=
\mathbf{\hat H}\,\mathrm{diag}(\mathbf\Theta_1,\dots,\mathbf\Theta_L)\,\mathbf{\bar H}.
\]
Under the worst-case LoS setting where each subchannel has rank \(1\), the aggregate channel can satisfy
\[
\operatorname{rank}\!\left(\sum_{i=1}^L \mathbf{\hat H}_i\right)=L,\qquad
\operatorname{rank}\!\left(\sum_{i=1}^L \mathbf{\bar H}_i\right)=L,
\]
so the effective rank can grow with \(L\) [2505.01170]. This is the paper’s formal argument for why physically separated RISs improve AirFC fidelity, especially in LoS-dominated environments.

The empirical evaluations use MNIST and Fashion-MNIST. In the non-trainable setting, increasing the total number of reflecting elements \(M\) reduces imitation error, and increasing the number of RISs \(L\) also reduces imitation error. One reported example is that at \(M=100\), the imitation error is about \(2.8\) for \(L=5\) versus about \(4.6\) for \(L=1\), corresponding to roughly a \(39.1\%\) reduction. Another reported example is that at \(K=30\) dB, the imitation error is about \(20.8\) for \(L=5\) versus \(78.3\) for \(L=1\), corresponding to a \(73.4\%\) reduction [2505.01170].

The trainable setting shows related trends. Increasing \(P_{\max}\) and increasing \(M\) improve classification accuracy; in the non-trainable case, increasing the Rician factor \(K\) worsens emulation error because the channel becomes more LoS-dominated and rank-deficient; in the trainable case, larger \(M\) improves performance through higher passive beamforming gain and better noise suppression; and multiple RISs significantly improve both emulation error and classification accuracy, with the strongest gains in LoS-dominated scenarios [2508.01840]. The best classification accuracy is reported for the trainable \(|v_i|\le 1\) case, whereas the practical model uses \(|v_i|=1\) [2508.01840].

## 6. Assumptions, architectural boundaries, and practical limitations

The main AirFC models assume that the direct Tx-Rx link is blocked, that the number of transmit and receive antennas equals the FC-layer dimension \(N\), that RIS elements are passive and impose unit-modulus phase shifts, and that higher-order reflections are ignored [2505.01170]. They also assume that the channel is known or estimated well enough to optimize the over-the-air parameters, which makes CSI acquisition and calibration a practical concern [2505.01170].

The framework is directed at **linear FC layers**. Non-linear layers remain digital or are approximated by other analog blocks [2505.01170]. The performance is channel-sensitive, particularly to rank structure and to the LoS/NLoS balance, and physical deployment cost increases with multiple RISs even though multi-RIS configurations improve fidelity [2505.01170]. The later trainable formulation reduces dependence on explicit CSI through reciprocity-based distributed learning, but this comes at the cost of noisy gradient estimation, especially when \(P_{\max}\) is small [2508.01840].

These assumptions delimit what AirFC currently means in the literature: it is a wireless realization of FC-layer linear algebra, not a complete analog replacement for arbitrary neural-network architectures.

## 7. Terminological scope and adjacent research lines

The term **AirFC** is specific to FC-layer realization over the air in RIS-assisted MIMO/OAC systems [2505.01170]. It is distinct from several adjacent lines of work that also combine wireless propagation and computation.

One nearby direction is **fluid-antenna-enhanced over-the-air computation**, where antenna positions or antenna-port selections are optimized to reduce aggregation MSE for sums of user symbols rather than to emulate an FC matrix. In that setting, the design variables are transmit coefficients, a receive decoding vector, and an antenna position vector, and the objective is AirComp MSE minimization under antenna-placement constraints [2312.15244]. A later copula-based analysis studies FA-assisted AirComp under spatially correlated fading, derives a closed-form CDF of the aggregation MSE, and shows that FA deployment reduces MSE relative to fixed-antenna systems, with gains that diminish as correlation strengthens [2605.03280].

Another adjacent direction is **AirFL**, including CSI-free non-coherent over-the-air federated learning, where the task is aggregation of model updates rather than FC-layer emulation. NCAirFL, for example, uses square-law non-coherent detection, binary dithering, and memory-based error compensation to obtain an \(\mathcal O(1/\sqrt{T})\) convergence rate for smooth non-convex objectives without instantaneous CSI [2411.13000]. By contrast, AirFC targets the realization of a neural-network FC layer itself.

A frequent misconception is to treat AirFC as a generic label for any over-the-air learning or federated-learning method. That usage is not supported uniformly across the cited literature. In particular, the paper on **Anarchic Federated Learning** explicitly states that the string “AirFC” does not appear there and that the correct concepts are AFL, AFA-CD, and AFA-CS [2108.09875]. Within the current arXiv literature represented here, the precise technical meaning of AirFC is therefore the RIS-assisted over-the-air implementation of fully connected neural-network layers rather than fluid-antenna AirComp or asynchronous federated learning.

Source: https://www.emergentmind.com/topics/airfc