---
title: Large-Scale Optoelectronic Neurons
url: https://www.emergentmind.com/topics/large-scale-optoelectronic-neurons-oens
type: topic
---

# Large-Scale Optoelectronic Neurons

Large-scale optoelectronic neurons (OENs) are hybrid neural-processing elements in which optics carries communication, broadcast, weighting, interference, diffraction, or summation, while electronics provides photodetection, thresholding, nonlinear activation, state storage, calibration, or control. In the cited literature, the term covers several distinct but related abstractions: wavelength-division multiplexed perceptrons driven by Kerr microcombs, multi-operand interferometric dot-product engines, diffractive detector-pixel neurons, coherent VCSEL-based neurons with inline homodyne nonlinearity, superconducting loop neurons that communicate optically and compute electronically, sensor-array MAC pixels, and spiking electro-photonic tiles [2101.12356] [2305.19592] [2008.11659] [1805.01947].

## 1. Scope of the term and major architectural classes

The literature does not use *large-scale optoelectronic neuron* in a single narrow sense. In microcomb and interferometric ONNs, an OEN is usually a perceptron-like unit that computes a weighted sum optically and applies its activation electronically or through device transfer functions. In diffractive and sensor-array systems, the “neurons” may be detector pixels or demodulator pixels whose optical field summation is followed by electrical readout. In spiking neuromorphic systems, the neuron is an excitable electronic or superconducting circuit that emits optical spikes for communication while retaining synaptic and dendritic computation in electronics [2101.12356] [2008.11659] [1809.02572] [2106.14803].

This plurality is not incidental. It reflects different answers to the same systems problem: how to combine optical fan-out, optical bandwidth density, and optical parallelism with electronic thresholding, memory, and control. Some platforms pursue dense matrix algebra for feedforward deep learning; others prioritize asynchronous spike routing, refractory dynamics, or online plasticity. A plausible implication is that “large-scale” in this field denotes not only physical neuron count, but also scalable fan-in, fan-out, synapse virtualization, and routable interconnect.

| Architecture family | Optics–electronics partition | Representative scale/result |
|---|---|---|
| Kerr-microcomb perceptron | Optical weighting and summation; electronic activation | 49 wavelengths, 11.9 Giga-OPS/s, 95.2 Gbps |
| MOMZI/MOON | Multi-operand interferometric dot product; electronic readout/activation | 85.89% SVHN, 128×128 PTC analysis |
| Diffractive DPU | Optical propagation and modulation; detector-pixel activation | Millions of neurons, up to 270.5 TOPs/s |
| Coherent VCSEL neurons | Optical MAC and inline homodyne nonlinearity | 7 fJ/OP, 25 TeraOP/(mm²·s) |
| Photonic neural field / CIS OEN | Optical mixing or demodulation; electronic readout and control | 1.3 P MAC/s single wavelength / 12.6 POPS |
| Superconducting and spiking OENs | Optical communication; electronic or superconducting computation | Up to \(10^6\) neurons per wafer in scaling studies |

The table organizes the principal families, but the category boundaries remain porous. For example, the microcomb perceptron is an OEN because photodetection and activation remain electronic; the diffractive DPU is also treated as an OEN because sCMOS pixels act as perceptron-like optoelectronic neurons; and superconducting systems are OENs because photons carry spikes while Josephson and loop circuits implement synaptic and neuronal functions [2003.01347] [1805.01942].

## 2. Core computation primitives and neuron models

A recurrent formal template is the perceptron equation
\[
y = \phi\!\left(\sum_{i=1}^{N} w_i x_i + b\right),
\]
implemented by optical broadcast and weighting, followed by photodetection and activation. In the soliton-crystal Kerr-microcomb perceptron, synapses are mapped onto comb wavelengths, a programmable waveshaper sets per-wavelength attenuation, the input vector is time-division multiplexed and broadcast to all wavelengths by an electro-optic modulator, a dispersive delay aligns the diagonal terms, and the photodiode sums the aligned optical powers. The photodiode current obeys \(I(t)=R\sum_{i=1}^{N}P_i(t)\), and the activation was a sigmoid applied offline in the demonstration; the paper also notes hardware activation through a Mach–Zehnder modulator or amplifier saturation [2003.01347].

In multi-operand interferometric neurons, the primitive is not wavelength-time broadcast but coherent aggregation inside a single interferometer. A multi-operand optical neuron partitions the modulation region into \(K\) independently driven operands, and a multi-operand Mach–Zehnder interferometer (MOMZI) realizes a \(K\)-term weighted sum in one device. The reported dual-arm transfer takes the form
\[
y=\cos^2\!\Big(\frac{1}{2}\big(\sum_{i=1}^{K}\theta_{u,i}-\sum_{i=1}^{K}\theta_{l,i}\big)+\phi_b\Big),
\]
so weighting and signed accumulation occur inside the interferometric transfer, while photodiodes, TIAs, and optional ADCs provide readout and downstream activation [2305.19592].

Other OEN families relocate the primitive to different physical observables. In the reconfigurable diffractive processing unit, each detector pixel acts as a perceptron-like optoelectronic neuron: free-space propagation performs linear transforms, and photodetection implements the nonlinear activation \(I(x,y)=|E(x,y)|^2\). In the CMOS image-sensor OEN, a lock-in demodulator pixel performs a signed multiply by complementary charge accumulation, with
\[
(Q_1+Q_3)-(Q_2+Q_4)=(2C-1)(2R-1),
\]
and summation over the illumination window yields a dot product digitized by 8-bit ADCs [2008.11659] [2511.04136].

Spiking OENs use a different neuron model. Superconducting optoelectronic loop neurons store state as flux in synaptic integration and neuronal integration loops, trigger when the integrated current exceeds a Josephson threshold, and emit optical spikes through an amplifier chain driving a light source. Their phenomenological reduction leads to a nonlinear leaky-integrator ordinary differential equation for each dendrite,
\[
\beta \frac{ds}{d\tau}=r(\phi,s;i_b)-\alpha s,
\]
with separate refractory feedback for the soma [2210.09976]. This places OENs in direct contact with both ANN-style weighted-sum neurons and event-driven excitable spiking neurons.

## 3. Wavelength, interferometric, and coherent-array OENs

The soliton-crystal Kerr-microcomb perceptron is a canonical large-scale OEN in the ONN sense. It maps 49 synapses onto 49 wavelengths with spacing \(\Delta f \approx 48.9\) GHz, uses \(N=49\) symbols at 11.9 Gbaud with \(t \approx 84\) ps, and reports 11.9 GFLOPS at 8 bits per FLOP, corresponding to 95.2 Gbps. The same work reports 93.75% experimental accuracy for handwritten digit recognition in binary pairs such as 0 vs 6, 86.67% for cancer-cell classification, OSNR \(>28\) dB, waveshaper attenuation range of 35 dB, latency of \(\sim 64\,\mu\text{s}\) dominated by the \(\sim 13\) km fiber spool, and a scaling path in which wavelength, time, and spatial multiplexing support deep optoelectronic neural networks with one microcomb provisioning many synapses [2003.01347].

The MOMZI/MOON line addresses a different bottleneck: the area and loss cost of single-operand MZI meshes. By placing \(K\) operands inside one interferometer and parallelizing MOMZIs per row, the architecture eliminates \(O(n)\) cascades per optical path. Experimentally, the reported 4-op MOMZI chip achieves 85.89% measured accuracy on SVHN with 4-bit voltage control precision. At the architecture-analysis level, a 128×128 MOMZI-based photonic tensor core is reported to have \(\sim 49\times\) lower propagation delay, \(\sim 256.7\) dB lower propagation loss, \(\sim 6.2\times\) smaller total device area when using 128-op MOMZIs, and 127× fewer high-speed MZI modulators than single-operand counterparts, while preserving comparable matrix expressivity through device-aware training [2305.19592].

Coherent VCSEL neural networks push OENs toward a different operating point: ultralow electro-optic drive and inline nonlinearity. Injection-locked VCSEL arrays encode inputs and weights as optical fields, diffractive fanout provides parallel channels, and balanced homodyne detection performs photoelectric multiplication while also supplying an instantaneous phase-sensitive nonlinearity. The reported system reaches 7 fJ/OP full-system energy efficiency, 25 TeraOP/(mm²·s) compute density, measured modulation power \(P_m=V_\pi^2/R_{\text{VCSEL}}=3.6\) nW with \(V_\pi=4\) mV and \(R_{\text{VCSEL}}=4.3\,\text{k}\Omega\), and hardware inference accuracy of \(93.1 \pm 2\%\) on MNIST over 1000 test images, matching 98% of the model’s simulated accuracy of 95.1% [2207.05329].

These three families illustrate a central divide within large-scale OEN design. Microcombs scale synapses spectrally; MOMZIs collapse multiple operands into a single interferometer; coherent VCSEL arrays reduce electro-optic energy and embed nonlinearity at the detector. All three depend on photodetection as the interface between optical parallelism and electronic state evolution, but they distribute complexity differently across source engineering, filter calibration, interference control, and readout electronics.

## 4. Spatial-field, diffractive, and sensor-array OENs

Spatially distributed OENs replace explicit per-synapse optical routing with field propagation, detector arrays, or time-domain pixel computation. The reconfigurable diffractive processing unit is the clearest example. It implements various diffractive feedforward and recurrent neural networks by combining a DMD for input coding, an 8-bit phase SLM for trainable diffractive modulation, and an sCMOS detector whose pixels provide optical summation plus complex activation through photodetection. The platform supports “millions of neurons,” operates at 56 fps for D2NN inference and \(\sim 70\) fps for D-RNN read-in, reaches 133.4 TOPs/s for D2NN and D-NIN and 270.5 TOPs/s for D-RNN, and reports system energy efficiencies of 2.889 TOPs/J for D2NN and 5.855 TOPs/J for D-RNN. A key contribution is adaptive training: direct transfer of a three-layer D2NN to hardware yielded 63.9% MNIST test accuracy, while measured-field adaptive training improved this to 96.0% using the full training set and 93.9% with a 2% mini-set [2008.11659].

The photonic neural field on silicon takes a still more distributed view. Here the “neurons” are virtual samples of a continuous speckle field generated by multimode interference in a 25 \(\mu\)m wide, 39 mm long silicon spiral waveguide with footprint \(\approx 2.4 \times 2.0\) mm². The prototype uses \(N=65\) spatial points and \(K=4\) time samples per symbol, giving \(NK=260\) virtual neurons from a single wavelength, and with \(L=5\) wavelengths reaches \(NKL=1300\) virtual neurons. Using the throughput expression \([6NM+2(M-1)]K/\tau\), the work reports \(1.333\times10^{15}\) MAC/s \(\approx 1.3\) P MAC/s for a single input wavelength. For chaotic time-series prediction at 12.5 GS/s, it reports NMSE \(\approx 0.039\) with a single wavelength and 0.018 with five wavelengths [2105.10672].

Sensor-like OEN arrays appear in two distinct forms. One is the transparent optoelectronic neuron array that couples transparent 2D MoS\(_2\) phototransistors to twisted-nematic liquid-crystal modulators. A 100×100 transparent array on a 1 cm × 1 cm substrate functions as a self-modulating nonlinear filter for incoherent broadband light. Under glare-reduction conditions, glare transmission was reduced by 74% relative to \(V_{dd}=0\) V while the non-glare region dropped by only \(\sim 9\%\), and the fabricated array had 98.94% functional yield [2304.13298]. The other is the CMOS image-sensor OEN for transformer inference, where a 2048 × 3072 array of demodulator pixels, hybrid-bonded to mixed-signal electronics and HBM, is analyzed for GPT-3 inference. With all required optoelectronic devices and circuits integrated in a chiplet about 2 cm by 3 cm, the reported figures are 12.6 POPS, 74 TOPS/W, and 19 TOPS/mm² for 175 billion parameters using a 40 nm CMOS process node [2511.04136].

What unifies these otherwise dissimilar systems is the replacement of explicit neuron-by-neuron photonic routing with field-level or array-level optical computation. This suggests a second major branch of the OEN literature: not optical neurons as individual photonic devices, but optoelectronic neuron fabrics in which pixels, samples, or demodulation sites instantiate the computational graph.

## 5. Spiking, event-driven, and superconducting OENs

The spiking branch of large-scale OEN research is dominated by architectures that reserve optics for communication and keep synaptic, dendritic, and somatic dynamics electronic. Superconducting optoelectronic loop neurons use superconducting single-photon detectors, Josephson junctions, storage loops, microscale LEDs, and multi-planar dielectric waveguides. In one formulation, optical communication budgets assume \(\sim 10\) photons delivered per synapse to cover routing loss and support both firing and update taps, synaptic firing can be triggered with one photon, the electrical energy required to generate an optical burst of \(10^4\) photons at \(\eta_{\text{LED}}=10^{-3}\) is \(\sim 1.6\) pJ per optical spike, and with full amplifier-chain efficiency \(\eta_{\text{amp}}\approx10^{-4}\) it is \(\sim 16\) pJ per spike. Scaling analyses report \(\sim 8100\) neurons on 1 cm × 1 cm with \(\sim 1\) mW device power, and \(\sim 10^6\) neurons with \(\sim 2\times10^8\) synapses on a 300 mm wafer with \(\sim 1\) W device power and coherent oscillations over large-data-center area at 1 MHz [1805.01947] [1805.01942].

The broader “optoelectronic intelligence” framework argues that the largest cognitive systems will use photons for communication and Josephson circuits for computation, with operation at 4 K enabling single-photon detection and silicon light sources. It gives a direct scaling rule for coherent integration, \(D_{\max}\approx c/(nf)\), and uses it to argue that a large-data-center area of \(10^5\,\text{m}^2\) can be integrated at 1 MHz and that Earth-scale integration is possible at theta-band frequencies of 4 Hz [2010.08690] [1809.02572]. A complementary design study compares semiconductor and superconducting OEN platforms, emphasizing that semiconductor receivers require roughly 1000× more optical power than superconducting receivers for identical links, while superconducting systems must solve source driving, serial biasing, and cryogenic integration [2106.14803].

Semiconductor spiking OENs form a separate line. The laser spiking neuron in a photonic integrated circuit combines a balanced photodetector pair, a two-section DFB laser, and an SOA in a broadcast-and-weight WDM architecture. It demonstrates simultaneous excitation, inhibition, and summation across eight wavelength channels, output linewidth \(\Delta\lambda_{\text{out}}<0.001\) nm, regenerated spike widths of 0.2–0.3 ns, and closed-loop gain with \(\approx 3\) dB margin at \(I_{\text{SOA}}\ge105\) mA, implying a potential upper-bound spike rate of \(\approx 5\) GHz from a refractory interval of \(\approx 200\) ps [2012.08516]. The RTD–photodetector–VCSEL artificial spiking neuron instead uses resonant tunnelling diode excitability and produces \(\sim 100\) ns optical spiking responses with refractory period \(T_{\text{ref}}\approx90\) ns, while theory for a monolithic nanoscale implementation indicates reliable triggering of two spikes at 300 ps separation, corresponding to \(\sim 3.3\) GHz [2206.11044].

SEPhIA extends the spiking OEN concept by making laser count itself a scaling variable. One multi-wavelength source is shared across many spiking neurons in an optical tile, so each neuron “owns” a wavelength channel and modulates it with a compact microring resonator modulator. For \(N_T=16\), the lasers-per-neuron ratio is \(1/16\), the multi-layer optoelectronic SNN reaches 91.35% test accuracy on a four-class spike-encoded dataset, and the reported energy is \(\approx 2.516\) pJ/spike for a full neuron path at \(N_T=16\) and \(\approx 44\) fJ/spike for a minimal interlink [2510.07427]. In these architectures, “large-scale” is inseparable from optical fan-out, wavelength reuse, and event-driven sparsity.

## 6. Scaling laws, programmability, and recurring limitations

Large-scale OEN papers are unusually explicit about scaling laws. In the microcomb broadcast-and-delay perceptron, per-neuron throughput obeys
\[
\text{GOPS}=\frac{2N}{2N-1}R,
\]
which approaches \(2R\) for large \(N\), and the architecture also states the layer-capacity condition \(M^{(k-1)}M^{(k)}\le N_{\text{comb}}\) when a single comb provisions all synapses of a fully connected layer [2101.12356]. In MOMZI-based photonic tensor cores, delay and insertion loss remain essentially one-interferometer-per-path quantities rather than growing with cascade depth, which explains the reported \(\sim 49\times\) delay reduction and \(\sim 256.7\) dB propagation-loss reduction at 128×128 [2305.19592]. In the CIS OEN, throughput scales as
\[
Y = \frac{2 f_{\text{clk}} C_T C_W}{r},
\]
making large pixel arrays and DAC sharing central to the reported 12.6 POPS [2511.04136]. In the photonic neural field, throughput follows \([6NM+2(M-1)]K/\tau\), yielding \(1.3\) P MAC/s on a few-mm² footprint [2105.10672].

Programming strategies are equally diverse. Microcomb perceptrons load offline-trained weights into a commercial waveshaper and note that in situ training is feasible by iterative adjustment of attenuation values [2003.01347]. MOMZI systems calibrate each segment’s phase–voltage curve, fit the \(\cos^2\)-based transfer, and use hardware-aware training with process variation, thermal crosstalk, quantization, and dynamic noise in the loop [2305.19592]. The diffractive DPU introduced measured-data-driven adaptive training, in which experimentally measured intermediate fields are fed back into downstream retraining [2008.11659]. CIS OENs rely on INT8 PTQ or QAT with LLM.int8-style outlier handling, and their noise analyses indicate minimal impact from quantization formats and hardware-induced errors under the studied conditions [2511.04136]. Superconducting loop-neuron work, by contrast, often emphasizes physically native plasticity, including binary or multi-level flux-quantum memories and STDP triggered by single photons [1805.01947].

The principal limitations recur across otherwise dissimilar platforms. Optical-power budgets constrain fan-out and tile size in WDM systems; microcomb architectures note OSNR, comb flatness, photodiode saturation, modulator bandwidth, and calibration overhead; coherent VCSEL neurons identify phase stability, detector linearity, and ADC/DAC overhead; diffractive systems identify alignment, aberrations, refresh rates, and precision limits; transparent TPT–LC arrays are limited by LC speed, polarization dependence, and uniformity; semiconductor OEN roadmaps emphasize the challenge of integrating III–V sources and ultra-low-capacitance photodiodes; superconducting OENs inherit cryogenic infrastructure, source-driving, and serial-bias constraints [2101.12356] [2207.05329] [2008.11659] [2304.13298] [2106.14803].

A recurring source of ambiguity is that throughput, energy, and neuron-count figures are not directly comparable across subfields. Some papers count multiply and accumulate separately as two operations; others report MAC/s, TOPS/W, or system energy efficiency with different accounting boundaries; still others treat detector pixels or virtual samples as neurons. This suggests that the field’s central unifying question is not the identity of a single “best” neuron device, but how optics and electronics should be partitioned for a target regime of fan-out, precision, latency, programmability, and physical scale. Across the current literature, large-scale OENs remain a family of architectures rather than a settled canonical design.

Source: https://www.emergentmind.com/topics/large-scale-optoelectronic-neurons-oens