- The paper develops a physics-consistent model that separately quantifies encoding nonlinearity, mutual-coupling-induced structural nonlinearity, and network depth in programmable-metasurface neural networks.
- Phase encoding combined with strong mutual coupling enables near-perfect single-layer regression, while weak coupling increases NMSE by more than three orders of magnitude at moderate task difficulty.
- Depth improves expressivity even with shared weights, but purely multiplicative linear encoding remains ineffective for small inputs unless sufficient depth or affine encoding is introduced.
Motivation and problem statement
Programmable wave-domain computing (pWDC) seeks to offload specific-purpose computations from digital processors to the wave domain, exploiting reconfigurable wave–matter interactions for gains in latency, energy, and footprint. A central obstacle is that wave-based systems at moderate signal levels implement linear input–output mappings, whereas useful computation—particularly neural-network-style function approximation—requires non-linearity. Prior approaches inject non-linearity through RF-digital conversions, non-linear circuit blocks, or digital backends, all of which assume input data encoded into the input wavefront.
This paper (2603.13602) studies an alternative: structural input encoding, in which input data is written into the scattering properties of tunable elements of a programmable metasurface (PM). Under this scheme two distinct non-linear mappings arise: the encoding non-linearity from control variables to load reflection coefficients, and the structural non-linearity from those reflection coefficients to the output wavefront, induced by mutual coupling (MC) between PM elements. The authors identify a gap in prior work: existing WPNN demonstrations either conflate these mechanisms or cannot toggle them independently, and none systematically examines their interplay with network depth. The paper's stated contribution is the first systematic, physics-consistent study of how encoding choice, MC strength, and depth jointly govern expressivity in PM-based wave-based physical neural networks (WPNNs).
Physics-consistent system model
The system is modeled with multiport network theory (MNT). The cavity comprises NA​ antenna ports and NS​ tunable lumped loads embedded in static scattering objects described by a scattering matrix S. The end-to-end channel matrix is
H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,
with Φ=diag(r). Passivity guarantees spectral radius below unity, admitting a Neumann-series expansion whose higher-order terms are precisely the multi-bounce paths responsible for structural non-linearity. When SSS​=0 (no MC), H depends affinely on r; when additionally SRT​=0, the dependence is linear. The encoding function c maps real control variables to complex reflection coefficients; its non-linearity is a separate, independently controllable factor.
Notably, this is the first model-based analysis of multi-layer WPNNs with structural input encoding whose parameters derive from a full-wave solver; prior work was either model-agnostic, neglected MC, or used phenomenological coupled-mode parameters untied to any concrete implementation.
The physical platform is a compact D-band rich-scattering cavity (NS​0 mm) repurposed from a wireless network-on-chip design: a stratified dielectric stack (Rogers RO3003, Si, AlN) enclosed by conductive walls, with a NS​1 array of patch-based PM elements on top and 10 irregularly placed antennas below. A single ANSYS HFSS simulation yields the full NS​2 scattering matrix over 481 frequency points spanning 110–170 GHz, with operation at 140 GHz.
Expressivity is probed by scalar regression of filtered Gaussian noise sampled at 600 points in NS​3, with a 4th-order Butterworth cut-off NS​4 controlling target difficulty. Targets are standardized and scaled to lie within the physically reachable readout range; 200 samples train and 400 test each configuration, with normalized mean squared error (NMSE) as the metric, medianed over 50 random target functions.
Architectures and controlled factors
Two multilayer architectures are compared. In both, trainable weights enter through the control vector of each layer's PM-parametrized cavity, layers are cascaded unidirectionally via isolators, and the scalar readout is the mean real part (I-channel) across the five receive ports. In the independent-weights architecture each layer has its own weights; in the shared-weights architecture all layers reuse one weight vector. The shared-weights variant is the key methodological device: it isolates the effect of depth from the effect of additional trainable parameters.
Three factors are controlled systematically:
| Factor |
Mechanism |
Levels |
| Encoding non-linearity |
Choice of encoding function |
Phase encoding (NS​5, strongly non-linear) vs. linear encoding (NS​6, affine) |
| Structural non-linearity |
Time-gating truncation time NS​7 as MC proxy |
NS​8 |
| Depth |
Number of layers NS​9 |
Varied per architecture |
Time gating is applied in post-processing on the wideband channel spectrum, suppressing multi-bounce paths beyond delay S0; this avoids re-simulating the structure for each MC level. Training is end-to-end gradient descent (Adam, PyTorch) through the fully differentiable forward model—a model-based in-silico strategy whose feasibility for experimental prototypes is supported by recent demonstrations of experimental MNT parameter estimation, though it remains an assumption here since no hardware is built.
Main results
Three findings stand out.
Phase encoding plus strong MC suffices in a single layer. With phase encoding and no time gating, the single-layer NMSE at medium difficulty (S1) is more than three orders of magnitude lower than at the weakest considered MC (S2 ns), and example approximations are visually flawless even at high difficulty (S3). Performance improves monotonically with S4. This implies that strong structural non-linearity enables shallow WPNNs—an attractive property experimentally, since each layer adds footprint, isolators, and cost.
Linear encoding fails fundamentally at small inputs. With linear encoding, NMSE remains high regardless of MC strength. The failure is structural: for small S5, all reflection coefficients S6 vanish, so the PM-dependent output contribution collapses and the readout becomes dominated by an uncontrollable, PM-independent term. Approximations are good only for larger S7. The authors note this could be mitigated by affine rather than purely multiplicative encoding—an open modification not tested.
Depth partially substitutes for structural non-linearity—and helps even without new parameters. At weak MC (S8 ns), going from one to two layers drops NMSE by roughly three orders of magnitude under phase encoding, with marginal further gains beyond. Critically, the shared-weights architecture captures most of this benefit: depth improves expressivity even when the trainable-parameter count is fixed. With strong MC, added layers still improve performance by nearly an order of magnitude (shared weights) to more than an order of magnitude (independent weights), though the single-layer result is already visually near-perfect. For linear encoding, substantial depth (S9) can alleviate the small-H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,0 shortcoming, but only for low-difficulty targets.
Analytical interpretation
The discussion section grounds these observations in limiting-case expansions of the readout. With phase encoding, the scalar readout is exactly the real part of a Fourier series in H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,1: in the strong-MC single-layer case an infinite series whose coefficient decay slows with stronger MC; in the no-MC multi-layer case a truncated series with maximal harmonic order equal to the depth H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,2. Thus depth generates Fourier terms up to order H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,3 even absent MC, explaining why stacking layers compensates weak coupling. A notable degenerate case: if H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,4, all intermediate coefficients vanish and expressivity collapses—so direct transmit-to-receive leakage is actually desirable in the no-MC multilayer regime. With linear encoding, the analogous expansions are power series in H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,5, again infinite under strong MC and degree-capped by H(r)=SRT​+SRS​(I−Φ(r)SSS​)−1Φ(r)SST​,6 without it. The authors caution that because the series coefficients are constrained by physics and parametrization rather than freely tunable, differences in basis orthogonality (Fourier vs. monomial) are likely not the dominant effect—a candid qualification of what the analysis can and cannot attribute.
Limitations and open questions
Several limitations are explicit. The study is entirely computational: no experimental validation is performed, and the idealized encoding functions (unconstrained continuous phase; ideal linear reflectance) do not correspond to realistic varactor loads, whose tuning traces discretized, non-trivial paths in the complex plane. The regression task is scalar; extension to matrix-valued functions where multivariate non-linearity matters is left open, as is the inclusion of a possibly non-linear decoding stage such as intensity-only quantized readout. Model-based training assumes an accurately estimable forward model, which the authors argue is plausible but unproven for this specific D-band platform. Finally, whether the observed depth–MC interchangeability persists for realistic encodings and multivariate targets remains unresolved.
Conclusion
This paper provides a controlled, physics-consistent decomposition of expressivity in PM-based WPNNs with structural input encoding, separating encoding non-linearity, MC-induced structural non-linearity, and depth. Its central quantitative results—that phase encoding with strong MC yields near-perfect single-layer regression, that depth recovers much of the lost expressivity when MC is weak even without additional trainable weights, and that purely multiplicative linear encoding has a fundamental small-input deficiency—translate directly into design guidance: prioritize platforms with strong inter-element coupling and carefully chosen encoding functions before investing in deeper stacks. Experimental realization on the proposed D-band cavity, with realistic tunable-element encodings, is the natural next step the paper leaves open.