Photonics-Based MAC Operations
- Photonics-based MAC operations are analog computing primitives that use optical modulation, interference, and photodetection to perform weighted summations.
- They enable high-speed, massively parallel computations for neural network inference, training, and wireless communications via WDM and integrated microring resonators.
- Despite their tens-of-GHz operation, challenges remain in laser power, thermal tuning, analog precision, and integration with electronic components.
Searching arXiv for recent and foundational papers on photonics-based MAC operations. Photonics-based multiply-accumulate (MAC) operations are analog or mixed-signal computing primitives in which optical carriers, optical interference, wavelength multiplexing, tunable resonant elements, and optoelectronic detection are used to realize weighted summation of input signals. In the literature represented here, the dominant formulation is the dot product or matrix-vector product, such as or , with multiplication implemented by optical modulation, filtering, or interference, and accumulation implemented by optical combining and/or photodetection (Mehrabian et al., 2018, Al-Qadasi et al., 2021). These MAC engines are central to photonic neural-network inference and training, convolution acceleration, wireless baseband processing, and more recent hybrid photonic-electronic matrix multiplication systems (Filipovich et al., 2021, Salmani et al., 2021, Gong et al., 14 Apr 2026). A recurrent theme across the field is that photonics offers tens-of-GHz operation, wavelength-division multiplexing (WDM), and massive parallelism, while practical deployments remain constrained by laser power, insertion loss, tuning overhead, detector sensitivity, DAC/ADC cost, and precision limits (Al-Qadasi et al., 2021).
1. Operational definition and computational model
In these systems, MAC denotes the weighted accumulation of input channels into one or more outputs. The canonical expressions appear repeatedly across the literature: the scalar inner product , the matrix-vector product , and the matrix product (Filipovich et al., 2021, Ji et al., 2024). In photonic implementations, one operand is typically encoded into optical intensities or multi-wavelength analog optical channels, while the other is encoded into programmable optical transfer functions, most commonly microring resonators (MRRs) or Mach-Zehnder interferometer/modulator (MZI/MZM) structures (Filipovich et al., 2021, Salmani et al., 2021, Al-Qadasi et al., 2021).
The defining physical distinction from digital MAC hardware is that accumulation is often obtained without an explicit electronic adder tree. In broadcast-and-weight architectures, wavelength channels are co-propagated on a shared waveguide and then summed by a photodetector, whose photocurrent is proportional to the total weighted optical power (Mehrabian et al., 2018, Salmani et al., 2021). In interferometric mesh architectures, multiplication is described as intrinsic to interference or modulation, while accumulation arises through optical combining and readout (Al-Qadasi et al., 2021). This suggests that the photonic substrate is best viewed as an analog linear-transform engine rather than as a direct optical analogue of a conventional binary arithmetic logic unit.
The computational role of these MACs is typically not arbitrary control-intensive logic. Instead, the cited work targets dense linear algebra and structured tensor operations: convolution in convolutional neural networks, gradient-vector computation in direct feedback alignment, matrix multiplication and approximate inversion in massive-MIMO, and general matrix multiplication in hybrid systems (Mehrabian et al., 2018, Filipovich et al., 2021, Salmani et al., 2021, Gong et al., 14 Apr 2026).
2. Broadcast-and-weight with microring resonators
A major architectural family is the silicon photonic broadcast-and-weight protocol, which uses WDM, microring modulators or resonators, and photodetection to implement dot products (Mehrabian et al., 2018, Salmani et al., 2021). Its high-level flow is consistent across applications: inputs are encoded on distinct wavelengths, multiplexed onto a shared waveguide, weighted by tunable microrings, and accumulated at the photodetector.
In the photonic CNN accelerator PCNNA, each input value is carried on a distinct wavelength, multiple wavelengths are multiplexed and broadcast, each channel is weighted by an MRR bank, and a photodiode sums the weighted optical powers into a single photocurrent (Mehrabian et al., 2018). The MAC relation is written as
The paper further states
and, when optical power is proportional to the electrical input,
This is the essential optical MAC: wavelength channels carry the vector entries, microrings encode programmable coefficients, and photodetection performs accumulation (Mehrabian et al., 2018).
The same logic is adapted for wireless communications in a silicon-photonic platform for massive-MIMO processing (Salmani et al., 2021). There, left-hand-side matrix elements are loaded into a modulation section using all-pass MRRs and encoded in the intensities of wavelength-multiplexed signals; right-hand-side matrix elements are encoded as weights using add-drop MRRs in the weight bank; the optical signals propagate through a shared waveguide; and a photodetector with a transimpedance amplifier performs the summation and electrical readout. The targeted workloads include matrix multiplication, matrix inversion via iterative approximations, and ZF/MMSE detection, with the receive model
and detection matrix
The paper explicitly notes that native broadcast-and-weight is limited to real and positive values, so preprocessing is required for negative and complex-valued quantities (Salmani et al., 2021).
A notable advantage of the broadcast-and-weight family is parallelism from WDM. Multiple input channels coexist in one waveguide without electrical time multiplexing, and multiple output dot products can be instantiated by replicating rows of weight banks (Mehrabian et al., 2018, Filipovich et al., 2021). However, the same papers make clear that this parallelism is conditional on spectral spacing, microring calibration, thermal stability, and the performance of DACs, ADCs, TIAs, and lasers (Salmani et al., 2021, Filipovich et al., 2021).
3. Microring arrays for neural-network inference and training
Microring-based MAC engines have been proposed both for inference and for in situ gradient computation during training. In the training architecture using direct feedback alignment (DFA), the photonic core is a CMOS-compatible silicon photonic system employing add-drop MRRs coupled to through and drop ports, with balanced photodetection used to realize signed weighting (Filipovich et al., 2021). The weight is expressed as
0
where 1 and 2 are the drop and through transmissions. With a bank of MRRs and WDM-encoded analog inputs, the photodetector computes
3
Extending to arrays, an 4 photonic weight bank computes 5 dot products in parallel: 6
The training-specific contribution is the use of DFA, where hidden-layer gradients are formed as
7
The photonic hardware computes the matrix-vector product 8, while the Hadamard product with 9 is performed electrically using TIAs with digitally controlled gain (Filipovich et al., 2021). The attraction of DFA for photonics is explicit: no sequential backward chain is required, and multiple layers can be updated in parallel.
Performance projections in that work are framed in terms of
0
counting one multiply and one add per MAC-style operation. The abstract states operation at trillions of MAC operations per second with less than one picojoule per MAC operation, and the representative design using a 1 photonic weight bank is projected at 20 TOPS, 1.0 pJ/op with thermally locked MRRs, 0.28 pJ/op with post-fabrication trimming, and 5.78 TOPS/mm2 (Filipovich et al., 2021). The experimental system is much less efficient because thermally tuned MRRs are slow, with an estimated 3 due to 170 4s tuning speed. The paper also reports multiplication experiments using 3900 random combinations of input and weight values with standard deviation of multiplication error 0.019 and effective resolution 6.72 bits for a single MRR; for a 5 array, the off-chip balanced-photodetector measurements had error standard deviation 0.098 and effective resolution 4.35 bits, and on-chip measurements had error standard deviation 0.202 and effective resolution 3.31 bits (Filipovich et al., 2021).
For CNN inference, PCNNA maps the receptive field and kernel to a flattened MAC: 6 Its central architectural observation is that convolutional connectivity is sparse and local. Without receptive-field filtering, the required ring count scales as
7
whereas with filtering it scales as
8
For AlexNet layer 1 with input 9 and 96 kernels of size 0, the paper reports about 5.2 billion microrings without filtering and about 35 thousand with filtering, a savings of more than 150,0001 (Mehrabian et al., 2018). It further claims potential speedup of up to 5 orders of magnitude for the optical core and more than 3 orders of magnitude for the full system, although the full-system limit is dominated by DRAM, DAC, and ADC overheads (Mehrabian et al., 2018).
These microring-based systems illustrate a central pattern of photonic MAC research: photonics executes the MAC-heavy linear transform, but memory movement, quantization interfaces, and tuning remain electronic bottlenecks.
4. Interferometric and mesh-based photonic MACs
A second major family uses MZMs, MZIs, and tunable beam-splitter meshes to implement matrix transformations (Al-Qadasi et al., 2021). In this formulation, a laser feeds an 2 mesh of tunable 3 nodes. The tunable beam splitter has transfer matrix
4
which simplifies to a parameterized 5 transformation in 6 and 7 (Al-Qadasi et al., 2021). A general matrix is decomposed as
8
Within this perspective, multiplication is “free” in the sense of being intrinsic to interference or modulation, while accumulation is produced by optical combining and photodetection (Al-Qadasi et al., 2021).
The same paper contrasts MZI/MZM meshes with MRR/MRM WDM systems. MZI/MZM architectures are broadband and flexible but large, lossy, and power-hungry because tuning is often thermo-optic. The rectangular 9 mesh is estimated to have optical depth 0 with 1 crossings (Al-Qadasi et al., 2021). By contrast, microring systems are smaller and generally lower capacitance, but are more temperature-sensitive and constrained by free spectral range and crosstalk.
The link-budget emphasis in this literature is crucial. For MZM systems, the received optical power depends on laser power minus fiber attenuation, coupling loss, silicon-waveguide attenuation, splitter loss, phase-shifter loss, directional-coupler loss, and implementation penalties. For MRR systems, the budget instead depends on in-band and out-of-band losses of modulators and weight rings, as well as splitter loss and penalties (Al-Qadasi et al., 2021). The same work gives a noise-limited relation between received optical power and achievable bit resolution,
2
to show that higher precision requires higher optical power or lower noise (Al-Qadasi et al., 2021). This is one of the clearest formal statements of a general limitation: photonic MAC scalability is tightly coupled to detector sensitivity, laser wall-plug efficiency, insertion loss, and electrical receiver overhead.
The paper directly compares such systems with 8-bit CMOS MACs in 28 nm, citing an effective operation energy of about 28.85 fJ/op for CMOS, and stating that MZM photonic MACs are about 3.53 to 17.54 worse and MRR photonic MACs about 2.65 to 136 worse than 8-bit CMOS across 1b–4b input resolutions (Al-Qadasi et al., 2021). The significance is not that photonics is slower—it is faster, with tens-of-GHz operation—but that system-level energy remains difficult once realistic I/O and optical-loss budgets are included.
5. Precision, signed arithmetic, and hybrid photonic-electronic reconstruction
A recurring misconception is that photonic MACs directly provide arbitrary-precision arithmetic. The cited work instead shows that precision is one of the central unresolved constraints. In the massive-MIMO broadcast-and-weight system, the photonic engine natively supports only real, positive values because encoding is intensity-based (Salmani et al., 2021). Complex arithmetic is handled by decomposing
7
and reducing
8
to four real-valued matrix multiplications (Salmani et al., 2021). Negative values are handled by shifting the left-hand-side matrix: 9 so that
0
This is not intrinsic signed arithmetic in optics; it is a preprocessing framework that makes intensity-only hardware usable for signed and complex workloads (Salmani et al., 2021). The same paper reports that 6-bit precision is insufficient at higher SNR, while 8-bit precision plus sign bit matches GPU performance in their study.
The most explicit precision-oriented response appears in the hybrid system "LightMat-HP" (Gong et al., 14 Apr 2026). There, photonics is used only for low-bit mantissa multiplication, while electronics performs accumulation, shifting, normalization, and result assembly. The system adopts block floating-point (BFP), selecting the shared exponent for a block as
1
and representing the block as
2
Its photonic core uses cascaded MZMs with intensity relations
3
and, in the linear region, 4 (Gong et al., 14 Apr 2026). High-precision mantissa multiplication is then reconstructed digitally from low-bit slice products. For example, if two 10-bit mantissas are split into two 5-bit slices each, the full product is reconstructed from
5
This design is explicit that photonic addition exists in principle, but chooses digital accumulation to avoid noise growth and improve scalability (Gong et al., 14 Apr 2026).
That paper reports that photonic multiplication accuracy is reliable only at low bit widths; on the prototype, 3-bit inputs showed low error, and error grows substantially as operand width increases. For BFP dot products with 4-bit or 5-bit slicing, relative error stayed below about 0.7%–0.75% in the reported experiments. Throughput scales from 209.25 GFLOPS to a peak of 6654.44 GFLOPS, energy efficiency from 1.23 GFLOPS/W at 6 to 39.14 GFLOPS/W at 7, and compared with BITLUME the system achieves up to 2.488 lower latency and about 1.489 better energy efficiency (Gong et al., 14 Apr 2026). A plausible implication is that hybridization is becoming a preferred route when arithmetic precision beyond low-bit analog capability is required.
6. Scalability, application mapping, and system bottlenecks
Across the cited work, scalability is never treated as a purely optical question. It is always jointly determined by waveguide bandwidth, WDM channel count, thermal tuning, device footprint, data movement, and the cost of crossing the photonic-electronic boundary.
For CNN inference, PCNNA exploits local receptive-field sparsity so that the number of required microrings scales with 0 rather than 1 (Mehrabian et al., 2018). The optical convolution time is modeled as
2
with
3
Its reported system organization separates a fast optical domain around 5 GHz from slower memory and interface domains, and it identifies DACs as the main bottleneck (Mehrabian et al., 2018).
For neural-network training with DFA, scalability is extended by a GeMM compiler that subdivides larger matrices into compatible blocks for the available hardware (Filipovich et al., 2021). The same paper notes that an optimized design with finesse 368 could support up to 108 channels, while still acknowledging constraints from laser power, detector capacitance, crosstalk, fabrication variation, and thermal drift.
For massive-MIMO, scalability is treated through matrix tiling and repeated chip use. The total runtime is modeled as
4
and
5
The architecture is constrained by the available number of wavelength channels 6 and MRRs 7, and the paper emphasizes that ADC, DAC, and TIA remain throughput bottlenecks (Salmani et al., 2021). It reports a propagation time of about 110 ps for one representative setting, but also notes throughput about 10 GS/s because components are mainly limited by ADC, DAC, and TIA.
For general matrix multiplication in LightMat-HP, scalability is addressed through tile-based scheduling. Input matrices
8
are partitioned into tiles, padded to
9
flattened, processed by the photonic processing unit, and digitally reconstructed (Gong et al., 14 Apr 2026). The paper’s large-scale simulation assumes 100 PPUs, 97 GS/s DACs/ADCs, 100 GHz modulators/detectors, 1 GHz system clock, and specific DRAM and SRAM bandwidths. This suggests that recent work increasingly treats photonic MAC as one stage in a system pipeline rather than as a standalone arithmetic core.
The scaling survey (Al-Qadasi et al., 2021) is more cautionary. It argues that current SOI silicon photonics is faster than CMOS digital MACs but usually less energy-efficient at the system level because laser inefficiency, tuning power, insertion loss, driver and receiver power, memory-interface overhead, and packaging all scale poorly. It further states that the dominant bounds arise from rated laser output power, the SNR required for a given input resolution, and the quadratic growth of tuning and control overhead.
7. Emerging directions and unsettled questions
The literature points to several distinct trajectories for photonics-based MAC operations. One is continued refinement of silicon photonic broadcast-and-weight systems using MRR arrays for neural inference, training, and communications (Mehrabian et al., 2018, Filipovich et al., 2021, Salmani et al., 2021). Another is the broader exploration of interferometric meshes and WDM-resonant systems under realistic link-budget and energy constraints (Al-Qadasi et al., 2021). A third is the move toward hybrid photonic-electronic arithmetic, where optics performs low-bit multiplication and electronics performs exact or precision-preserving accumulation and reconstruction (Gong et al., 14 Apr 2026).
There are also indications of more compact or alternative physical substrates. The abstract of "M3ICRO: Machine Learning-Enabled Compact Photonic Tensor Core based on PRogrammable Multi-Operand Multimode Interference" describes an ultra-compact photonic tensor core using programmable multi-operand multimode interference (MOMMI), a single-device programmable matrix unit, and a block unfolding method for near-universal real-valued linear transformations, with reported gains of 3.4–9.60 smaller footprint, 1.6–4.41 higher speed, 10.6–422 higher compute density, 3.7–123 higher system throughput, and superior noise robustness (Gu et al., 2023). However, the supplied text also states that the provided document contains only an empty LaTeX skeleton and does not expose the technical mechanism. The appropriate conclusion is therefore limited: the abstract asserts a photonic MAC-relevant tensor-core direction, but the mechanism cannot be reconstructed from the supplied paper text.
A related but non-photonic comparison point is the frequency-domain acoustic-wave MAC engine on lithium niobate, which uses second-order nonlinear acoustic response so that sum-frequency components encode products and all contributions landing in the same frequency bin are naturally accumulated (Ji et al., 2024). Its demonstrations of 4×4 and 16×16 matrix multiplication, 128×128 convolution, and a single-device footprint of 0.03 mm4 underline a broader trend: wave-based MAC hardware increasingly exploits multiplexed physical dimensions and nonlinear mixing to compute many products simultaneously. A plausible implication is that some concepts often associated with photonic tensor cores—synthetic frequency dimensions, natural accumulation via superposition, and single-device parallelism—are being generalized across wave platforms.
The principal controversy is not whether photonic MAC works physically; it is how much of the surrounding system can be made competitive with electronic accelerators. The surveyed work is consistent on this point. Photonics can provide tens-of-GHz operation, wavelength-level parallelism, and extremely low-latency analog linear transforms (Al-Qadasi et al., 2021). Yet practical systems remain bounded by signed and complex-value handling, analog noise, limited effective bit precision, thermal sensitivity, laser and coupling inefficiency, DAC/ADC/TIA overhead, and reconfiguration cost (Salmani et al., 2021, Filipovich et al., 2021, Gong et al., 14 Apr 2026). The field therefore increasingly frames photonic MAC not as a universal replacement for digital arithmetic, but as a specialized or hybrid substrate for MAC-dominated linear algebra where throughput, parallelism, and latency outweigh the penalties of analog interfaces and calibration.