---
title: Metasurface Diffractive Optical Networks
url: https://www.emergentmind.com/topics/metasurface-based-diffractive-optical-networks
type: topic
---

# Metasurface Diffractive Optical Networks

Searching arXiv for recent and foundational papers on metasurface-based diffractive optical networks.
Metasurface-based diffractive optical networks are coherent optical processors in which subwavelength metasurfaces serve as trainable diffractive layers, so that local optical modulation and free-space propagation jointly implement a computational mapping between an input and an output field-of-view. In this framework, metasurfaces composed of a two-dimensional array of millions of meta-units can realize precise control of optical wavefront with subwavelength resolution, and can therefore function as constitutive layers of optical neural networks, diffractive deep neural networks, spectropolarimetric encoders, or hybrid opto-electric front ends [2210.08369] [2007.12813] [2507.20229]. Reported realizations span object recognition, multi-task classification, permutation operations, wavefront sensing, spectropolarimetry, super-resolution direction-of-arrival estimation, and inverse-designed multifunctional optics from the visible to terahertz and microwave regimes.

## 1. Architectural foundations

The basic architectural element is a metasurface layer modeled as a thin transmissive mask with trainable local response. In the simplest phase-only form, a layer is written as $T(x,y)=e^{i\phi(x,y)}$ with unit amplitude, while polarization-dependent or vectorial designs use Jones-matrix descriptions or separate phase functions for different channels. A single-layer (“singlet”) optical neural network places one metasurface between an input plane and an output plane; multilayer variants use cascaded metasurfaces separated by free-space gaps, and examples include doublets, bilayers, three-layer stacks, and five-layer diffractive networks [2210.08369] [2506.18242] [2409.08423].

The physical realization depends on wavelength and multiplexing strategy. Near-infrared “smart glass” implementations used amorphous silicon nanopillars on silica with a square lattice of period $750$ nm and height $1\,\mu$m, together with isotropic and birefringent meta-unit libraries that provided $\sim2\pi$ phase modulation and polarization multiplexing [2210.08369]. Visible on-chip multiplexed diffractive neural networks used TiO$_2$ pillars on quartz with $p=400$ nm and $H=600$ nm, later bonded through a $100\,\mu$m OCA spacer to a CMOS sensor with $0.8\,\mu$m pixels [2107.07873]. Terahertz arrangeable diffractive neural networks used high-resistivity silicon nanofins at $0.291$ THz, with each diffractive layer implemented as an $80\times80$ array of nanofins over an aperture of approximately $4.12\,\text{cm}\times4.12\,\text{cm}$ [2506.18242]. Spectropolarimetric diffractive optical networks employed anisotropic Si nanobricks with periodicity $P=600$ nm and height $h=700$ nm, while spin-multiplexed nonlocal metasurfaces used crescent-shaped silicon nanopillars in a hexagonal lattice with period $P=1\,\mu$m and height $H=327$ nm [2507.20229] [2606.16938].

A recurring design pattern is phase-only encoding with subwavelength sampling. In the terahertz arrangeable network, pixel size $=515\,\mu$m ensured sub-wavelength sampling $(\lambda/2)$ across the layer [2506.18242]. In the visible on-chip platform, the single-channel neuron density was $(1/p^2)=6.25\times10^6\,\text{neurons/mm}^2$, and with two polarization channels the effective density became $2\times6.25\times10^6\,\text{neurons/mm}^2$ [2107.07873]. This suggests that metasurface implementations derive their practical expressive power not only from network depth but also from the very high areal density of trainable optical neurons.

## 2. Propagation physics and capacity

Forward propagation is generally formulated by alternating local metasurface modulation and scalar diffraction. A common expression is
$$
U_{\ell+1}(x,y)=\text{PSF}_{\ell\rightarrow \ell+1}\otimes\left[U_\ell(x,y)\,t_\ell(x,y)\right],
$$
or, in angular-spectrum form,
$$
U_{\ell+1}(x,y)=\mathcal{F}^{-1}\{\mathcal{F}[U_\ell\,t_\ell]\cdot \exp(jk_z\Delta z)\},
$$
with equivalent Rayleigh–Sommerfeld formulations also used extensively [2506.18242] [2210.08369]. For polarization-multiplexed visible metasurfaces, the local response can be written with a diagonal Jones matrix,
$$
J_{\text{meta}}(x,y)=
\begin{bmatrix}
T_{xx}(x,y)e^{i\phi_x(x,y)} & 0\\
0 & T_{yy}(x,y)e^{i\phi_y(x,y)}
\end{bmatrix},
$$
where $\phi_x$ and $\phi_y$ are independently designed for orthogonal polarizations [2107.07873].

Pancharatnam–Berry phase and related geometric-phase mechanisms are central in several metasurface diffractive networks. In the terahertz arrangeable network, rotating each nanofin by angle $\theta$ provided an abrupt geometric phase $\phi=2\theta$ covering $[0,2\pi]$ [2506.18242]. In the spin-multiplexed nonlocal metasurface, the cross-polarized channel used
$$
t_{RL}(x,y)=\sqrt{T_{RL}(0)}\,\exp[i\,2\theta(x,y)],
$$
while the co-polarized channel exploited the momentum-space transfer function
$$
H_{\text{edge}}(k_x,k_y)\simeq \alpha (k_x^2+k_y^2)/k_0^2,
$$
which approximates a Laplacian high-pass filter for optical edge detection [2606.16938]. In the spectropolarimeter, wavelength- and helicity-dependent phase control combined propagation phase and geometric phase through
$$
\Phi(x,y,\lambda,\text{pol})=\phi_p(x,y,\lambda)\pm 2\theta(x,y).
$$
[2507.20229]

The theoretical capacity of coherent diffractive networks has been analyzed in terms of the dimensionality of the implementable linear transformations. For a network with $K$ trainable surfaces, input field $\mathbf{x}\in\mathbb{C}^{N_{\rm in}}$, output field $\mathbf{y}\in\mathbb{C}^{N_{\rm out}}$, and layer neuron counts $\{N_k\}$, the achievable solution-space dimension is
$$
D(K)=\min\Bigl(N_{\rm in}N_{\rm out},\sum_{k=1}^K N_k-(K-1)\Bigr).
$$
If each layer has the same number of neurons, $N_k=N$, then
$$
D(K)=\min\bigl(N_{\rm in}N_{\rm out},KN-(K-1)\bigr).
$$
This linear growth with depth, մինչև saturation at the field-of-view limit, was connected to depth advantages in statistical inference, learning, and generalization [2007.12813]. A related scaling law appeared in all-optical permutation networks, where the permutation approximation error dropped precipitously once the total number of trainable meta-atoms satisfied $N\ge N_iN_o$; for a $400\times400$ permutation, roughly $N\gtrsim160$k was required [2206.10152].

## 3. Training and inverse design methodologies

Training is ordinarily performed by differentiating through the optical forward model. In object-recognition singlets, the forward pass starts from the complex amplitude of the coherently illuminated object, applies the metasurface phase mask, propagates to the output plane, and integrates intensities over class-specific detection zones. The class probabilities are then defined by normalized zone energies,
$$
p_c=\frac{I_c}{\sum_j I_j},
$$
and optimized with cross-entropy,
$$
L=-\sum_c t_c\ln p_c,
$$
optionally with an auxiliary contrast term to maximize inter-zone contrast and robustness [2210.08369]. The trainable variables are the per-meta-unit phases, updated with Adam after computing $\partial L/\partial \phi(x,y)$ by the chain rule through the diffraction integral [2210.08369].

Multi-task and fabrication-aware variants extend this basic recipe. The arrangeable diffractive neural network introduced a weighted multi-task loss
$$
L_{\text{multi}}(\{\phi_\ell\})=\sum_{i=1}^N w_i\,\mathcal{L}_i(U^{(i)}_{\text{out}},Y_i),
$$
and for two tasks
$$
L_{\text{multi}}=\mathcal{L}(U^{(1)}_{\text{out}},Y_1)+\beta\,\mathcal{L}(U^{(2)}_{\text{out}},Y_2),
$$
with $\beta\in\{0.5,1.0,1.5,2.0\}$ selected empirically to bias performance between MNIST and Fashion tasks [2506.18242]. Tri-channel wavelength-multiplexed classifiers used
$$
L_{\text{total}}=w_1L_{\text{MNIST}}+w_2L_{\text{FMNIST}}+w_3L_{\text{KMNIST}},
$$
with best-performing weights $(w_1,w_2,w_3)=(0.13,0.20,0.67)$ in the reported end-to-end optimization framework [2409.08423].

Several works moved beyond direct phase optimization. The phase correlation method for multiplexed metasurfaces replaced joint optimization of multiple channel phases with learned mappings such as $\phi_1(x,y)\simeq f_1[\phi_2(x,y)]$ and $\phi_3(x,y)\simeq f_2[\phi_2(x,y)]$, where small multilayer perceptrons reduced the multi-channel problem to a single-channel inverse design over the “master” phase profile [2412.13531]. The diffractive meta-neural network for super-resolution direction-of-arrival estimation pre-trained “mini-metanets” that mapped geometric parameters $(d_a,d_b,\theta)$ to multi-frequency Jones-matrix responses, and then backpropagated through both angular-spectrum propagation and the differentiable surrogate models [2509.05926]. Three-task multiplexed metasurfaces similarly used surrogate ANNs to predict $(|T|,\cos\phi,\sin\phi)$ from $(W_x,W_y)$, then trained geometry directly with Adam, initial learning rate $10^{-3}$, decay by $0.9$ every $10$k steps, batch size $32$, and approximately $2500$ epochs [2409.08423].

Robustness is commonly injected during training. Smart-glass classifiers included random perturbations to model nonuniform illumination, misalignment of object, metasurface, and detector, and distance fluctuations [2210.08369]. Diffractive permutation networks used a “vaccination” strategy, perturbing lateral shift, axial displacement, and in-plane rotation on each mini-batch, with
$$
\Delta x,\Delta y\sim\mathcal{U}(-0.672\lambda v,0.672\lambda v),\quad
\Delta z\sim\mathcal{U}(-2\lambda v,2\lambda v),\quad
\Delta\theta\sim\mathcal{U}(-4^\circ v,4^\circ v),
$$
to enforce misalignment tolerance [2206.10152]. Hybrid wavefront sensors and spectropolarimeters also injected detector noise, fabrication variability, quantization, and dynamic-range effects during end-to-end training [2602.16535] [2507.20229].

## 4. Multiplexing, reconfiguration, and multifunctionality

Metasurface-based diffractive optical networks use multiple physical degrees of freedom to exceed the single-task, single-channel behavior of conventional diffractive processors. Polarization multiplexing is among the most direct strategies. Near-infrared object-recognition smart glass employed birefringent meta-unit libraries with different phase shifts $\phi_x$ and $\phi_y$ for orthogonal polarizations, enabling one metasurface to route different recognition tasks to different outputs [2210.08369]. Visible on-chip multiplexed diffractive neural networks and bilayer multi-task classifiers further exploited polarization channels under the same wavelength, while maintaining phase-only operation at the network-design stage [2107.07873] [2409.08423].

Wavelength multiplexing extends this principle across spectral channels. Bilayer TiO$_2$ metasurfaces were used to encode up to three parallel tasks at $\lambda_1=450$ nm, $\lambda_2=550$ nm, and $\lambda_3=650$ nm under fixed $x$-polarization, with channel filtering at the detection plane [2409.08423]. The phase-correlation method likewise targeted dual-wavelength classification at $532$ nm and $633$ nm under the same polarization, converting a multi-wavelength phase-design problem into a reduced single-channel optimization [2412.13531]. Spectropolarimetric diffractive optical networks generalized multiplexing further by encoding both spectrum and Stokes parameters into a single-shot intensity fingerprint on a CMOS sensor [2507.20229].

Reconfiguration need not require active meta-atoms. The arrangeable diffractive neural network achieved task switching by reordering two pre-trained metasurface layers: Task 1 used the sequence $\{1\rightarrow2\}$ and Task 2 used $\{2\rightarrow1\}$, so that layer rearrangement alone changed the effective mapping without re-fabrication [2506.18242]. The diffractive magic cube network pushed mechanical reconfiguration much further by combining permutation, rotation, and translation of three cascaded phase-only metasurface layers. In that system, up to $6\times16\times29\approx4179$ non-independent mechanical states were available as “potential channels,” and the training objective jointly optimized a selected subset of channels for holography, focusing, or orbital angular momentum generation [2412.20693].

Spin multiplexing and hybrid preprocessing provide another route to multifunctionality. In the edge-enhanced nonlocal metasurface, the co-polarized channel performed real-time edge detection through momentum-space filtering, while the cross-polarized channel simultaneously executed single-layer diffractive classification through PB-phase modulation [2606.16938]. A plausible implication is that metasurface-based diffractive optical networks increasingly treat optical preprocessing, encoding, and inference as co-designed functions rather than isolated modules.

## 5. Demonstrated functions and reported performance

Representative demonstrations cover classification, sensing, linear transforms, and inverse-designed optical elements.

| System | Function | Reported result |
|---|---|---|
| Metasurface smart glass | 4-digit and 10-digit coherent object recognition | 4-digit: theoretical $99.14\%$, experimental $(N=116)$ $99.14\%$; 10-digit: theoretical $86.50\%$, experimental $(N=208)$ $78.37\%$ [2210.08369] |
| Polarization-multiplexed smart glass | Split-digit recognition; letter + style recognition | Experimental $90.99\%$ and $81.44\%$ for the two digit groups; $(N=168)$ $92.81\%$ letter identity and $100\%$ style [2210.08369] |
| On-chip visible multiplexed DNN | Simultaneous recognition of digital and fashionable items | Simulated accuracy $>99\%$ for both channels; experimental on-chip accuracy $\approx96\%$ [2107.07873] |
| Arrangeable THz DNN | Layer-reordered multi-task recognition | Simulation: MNIST $\approx93.1\%$, Fashion-MNIST $\approx87.2\%$; experiment: MNIST $ \sim75\%$, Fashion $ \sim70\%$ [2506.18242] |
| Bilayer multiplexed metasurface DNN | Dual-task and three-task classification | PM-DNN dual-channel: $97.72\%$ MNIST and $88.01\%$ FMNIST; tri-channel end-to-end joint: $96.48\%$ MNIST, $85.68\%$ FMNIST, $85.35\%$ KMNIST [2409.08423] |
| Edge-enhanced spin-multiplexed DNN | Single-layer classification with optical edge preprocessing | MNIST accuracy improved from $64.2\%$ to $80.7\%$ [2606.16938] |
| Hybrid metasurface-encoded OENN | Wavefront sensing for adaptive optics | At $N=3$ and $1\times1024$ MLP: mean SR $\simeq0.60$ versus $\simeq0.37$ for ANN-only; experiment: median SR $\simeq0.67$ versus $\simeq0.46$ [2602.16535] |
| DON spectropolarimeter | Spectrum and full-Stokes reconstruction | NIR: narrow-band MAE $=0.004$, broadband MAE $=0.031$, MRE $=7.39\%$, spectral resolution $6$ nm; chip-integrated visible: narrowband MAE $=0.007$, broadband MAE $=0.027$, MRE $=6.98\%$, spectral resolution $4$ nm [2507.20229] |
| DMNN | Super-resolution direction-of-arrival estimation | $0.5^\circ$ two-source resolution, mean absolute error $0.048^\circ$ for two incoherent targets within $\pm11.5^\circ$, AET $\approx1917$ [2509.05926] |
| MAODCNN | All-optical convolution plus diffractive decoding | MNIST: $86\%$ for $1$ conv $+$ $1$ diffractive, $\simeq94\%$ for $1$ conv $+$ $2$ diffractive; Fashion-MNIST: $79\%$ versus $73\%$ for DNN [2512.05558] |

Beyond classification and sensing, diffractive optical networks have been trained as general optical transforms. Permutation networks engineered through deep learning performed arbitrarily selected permutation operations and scaled to a $400\times400$ permutation, corresponding to $0.16$ million interconnects; a three-layer network was experimentally demonstrated at terahertz frequencies [2206.10152]. The diffractive magic cube network numerically and experimentally validated $144$-channel holograms, $108$-channel single-focus/multi-focus patterns, and $60$-channel OAM beam or comb generation, with mean experimental PCC or CC within $2\%$ of simulation [2412.20693].

Inverse-designed diffractive devices also reveal the continuity between optical neural networks and multifunctional flat optics. Two-layer adaptive deep diffractive neural networks selectively focused radiation over two well-separated spectral bands, achieved peak efficiencies $\eta_1,\eta_2>50\%$, surpassed the $\sim40\%$ bound of single-layer DOEs, produced Gaussian-shaped bandwidths as narrow as $\sigma=5$–$10$ nm, and generated super-oscillatory focal spots with $\alpha=0.4$ [2202.13518]. This suggests that metasurface-based diffractive optical networks should be understood not only as classifiers but also as a general inverse-design formalism for compact optical operators.

## 6. Fabrication, integration, applications, and limitations

Fabrication routes are strongly aligned with established nanofabrication workflows. Reported processes include a-Si deposition on SiO$_2$ followed by electron-beam lithography and reactive-ion etching, photolithography plus reactive-ion etching on $1$ mm Si wafers, TiO$_2$ nanofabrication by EBL, ALD, ion-beam etch, and resist stripping, and 3D printing of THz diffractive layers in UV-curable polymer [2210.08369] [2506.18242] [2107.07873] [2206.10152]. CMOS compatibility was emphasized repeatedly, and large-area manufacturing prospects included deep-UV lithography, nanoimprint lithography, and roll-to-roll replication [2210.08369]. Chip-level integration has already been demonstrated through bonding to commercial CMOS image sensors in both visible classification and spectropolarimetric systems [2107.07873] [2507.20229].

The attraction of these systems is their passive optical front end. For the smart-glass classifier, inference was reported as zero power consumption after the light source, with latency determined by light propagation time, approximately $10$ ps for $3$ mm, and “physics-guaranteed security” because no electronic read-out or digital image was ever created [2210.08369]. The terahertz arrangeable DNN similarly reported passive operation, zero additional electrical power beyond the THz source, and total latency below $100$ ps through both layers and free-space gaps [2506.18242]. In optical convolutional diffractive networks, numerical estimates placed total energy per inference at $O(10\,\text{pJ})$ and overall latency near $10$ ns when detector readout was included [2512.05558].

Applications follow directly from these properties. Reported scenarios include smart windows or eyewear, contact-lens cameras, embedded IoT sensors, security screening, nondestructive testing, portable biomedical terahertz spectroscopy, free-space optical communications, optical data storage and display, multi-focus microscopy, optical encryption, and onboard remote sensing from raw Sentinel-1 level-0 IQ data [2210.08369] [2506.18242] [2412.20693] [2503.13488]. The remote-sensing SIM-D$^2$NN reported performance levels around $90\%$ directly from real raw IQ data in terms of accuracy, precision, recall, and F1 Score, using four phase-only layers and two output antennas [2503.13488].

Several limitations recur across the literature. Point-by-point scanning output detectors constrain frame rate in terahertz demonstrations, and integrated focal-plane arrays are identified as necessary for real-time full-field readout [2506.18242]. Fixed phase patterns mean that truly dynamic reprogramming requires active meta-atoms or SLM integration [2506.18242]. Tight alignment tolerances remain important: reported values include feature alignment tolerance $<50$ nm over $500\,\mu$m fields, overlay tolerance $\lesssim50$ nm in visible on-chip systems, and lateral or axial sensitivity on the order of $\pm100\,\mu$m and $\pm500\,\mu$m in terahertz reconfigurable networks [2210.08369] [2107.07873] [2506.18242]. Scalar diffraction approximations can break down when inter-pillar coupling or polarization cross-talk grows, and metasurfaces are often inherently narrowband, while the capacity analysis in coherent diffractive networks assumes monochromatic coherence [2107.07873] [2007.12813].

A common misconception is that all metasurface-based diffractive optical networks are fully all-optical from input to decision. Some are: smart-glass classifiers and diffractive permutation networks operate as passive optical transforms whose outputs can be read directly as intensity patterns [2210.08369] [2206.10152]. Others are deliberately hybrid. Wavefront sensing used a metasurface encoder followed by a lightweight multilayer perceptron; the spectropolarimeter used a trained CNN decoder; the super-resolution DMNN used a lightweight electronic neural network after optical multiplexing [2602.16535] [2507.20229] [2509.05926]. The distinction is functional rather than conceptual: in all cases, the metasurface stack performs a learned optical encoding that shifts computation into wave propagation, while the electronic back end, when present, decodes a compressed optical representation.

Source: https://www.emergentmind.com/topics/metasurface-based-diffractive-optical-networks