Papers
Topics
Authors
Recent
Search
2000 character limit reached

All-Optical Systolic Arrays

Updated 14 July 2026
  • All-optical systolic arrays are photonic architectures where both matrix operands are encoded as dynamic optical pulse streams, enabling high-throughput computation.
  • They use an output-stationary design with time-staggered pulses and homodyne detection to perform precise multiply–accumulate operations.
  • Compact inverse-designed photonic modules and rigorous calibration strategies achieve dense computation while addressing challenges in scaling and integration.

Searching arXiv for the cited papers and closely related work on photonic/all-optical systolic arrays. {"query":"arXiv (Kim et al., 2024) photonic systolic array all-optical matrix-matrix multiplication", "max_results": 5} {"query":"arXiv (Lin et al., 15 Jul 2025) SystolicAttention fusing FlashAttention within a Single Systolic Array", "max_results": 5} {"query":"all-optical systolic array photonic arXiv Youngblood Rahimi Kari", "max_results": 10} All-optical systolic arrays are photonic matrix-computation architectures in which both multiplicand matrices are represented as optical signals and processed by a two-dimensional, heartbeat-like dataflow of multiply–accumulate cells. In the formulation numerically verified in "Photonic systolic array for all-optical matrix-matrix multiplication," the array is output-stationary: columns of A\mathbf{A} propagate vertically, columns of B\mathbf{B} propagate horizontally, and each processing element accumulates one entry of C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B} through homodyne detection of two traveling optical fields. The reported design uses adjoint-based inverse design to realize compact freeform crossings, branches, beam splitters, and grating couplers in 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m} footprints, and a simulated 4×44\times4 array yields a theoretical computation density of 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s} (Kim et al., 2024).

1. Conceptual basis and distinction from weight-stationary photonics

A systolic array is a regular grid of processing elements synchronized by a staged, wavefront-style movement of data. In electronic accelerators this organization is widely used for matrix multiplication, and the photonic analogue has often been sought in waveguide crossbars, Mach–Zehnder interferometer meshes, and inverse-designed linear networks. The crucial distinction drawn by the 2024 photonic proposal is that most earlier photonic accelerators are effectively weight-stationary systolic arrays: one matrix is encoded in propagating optical signals, while the other is embedded as static or slowly reconfigurable photonic parameters such as coupling ratios, phase shifts, or material distributions. That model corresponds to Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X} with W\mathbf{W} hard-coded or only slowly updated (Kim et al., 2024).

The all-optical systolic array removes that asymmetry. Both multiplicands are dynamic optical streams, and nothing in the photonic fabric acts as the matrix of weights. In the reported architecture, columns of A\mathbf{A} are injected from the bottom edge and move upward; columns of B\mathbf{B} are injected from the left edge and move rightward. Each cell indexed by B\mathbf{B}0 is associated with a fixed output location and accumulates the product terms for one matrix entry. This output-stationary organization is significant for workloads in which both operands change rapidly, especially transformer attention, where B\mathbf{B}1 and B\mathbf{B}2 are both data-dependent and must vary independently from token to token (Kim et al., 2024).

A common misconception is that any photonic matrix multiplier is already an all-optical systolic array. The reported design rejects that equivalence. It reserves the term for an architecture in which both matrices are encoded optically as pulse trains, the multiply–accumulate step is performed by optical interference and photodetection, and electronics are used only to integrate and read out photocurrent rather than to implement the linear algebra itself (Kim et al., 2024).

2. Optical encoding, systolic timing, and the homodyne MAC

The encoding scheme is symmetric in the two multiplicands. Matrix elements B\mathbf{B}3 and B\mathbf{B}4 are represented as real-valued amplitudes of optical pulses at carrier frequency B\mathbf{B}5, around B\mathbf{B}6, guided in a silicon slab waveguide in the TMB\mathbf{B}7 mode. Each element occupies a time slot, so each column becomes an amplitude-modulated pulse train. One stream is phase-shifted by B\mathbf{B}8, enabling homodyne detection of a real product at each cell (Kim et al., 2024).

The systolic schedule is defined by time-staggered injection. For the B\mathbf{B}9 outer-product example, C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}0 and C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}1 are injected simultaneously, followed by C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}2, then C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}3, then C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}4, with inter-pulse delays equal to the propagation time between adjacent cells. As the pulses move through the grid, successive diagonals satisfying C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}5 become active in turn. The resulting diagonal wavefront is the optical counterpart of the classic heartbeat of a systolic array. The array consequently exhibits fill, steady, and drain phases, and the outputs emerge in sequence from lower left to upper right (Kim et al., 2024).

At the cell level, the multiplication mechanism is coherent interference followed by balanced detection. For monochromatic inputs C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}6 and C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}7 incident on a C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}8 beam splitter, the intensity difference between the two output ports satisfies

C=ATB\mathbf{C}=\mathbf{A}^T\mathbf{B}9

for real amplitudes 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}0 and 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}1. For time-varying envelopes 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}2 and 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}3, the detected photon-count difference over a time window is

3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}4

After quantum-efficiency and charge conversion, the accumulated detector output is proportional to the temporal inner product of the two streams. The full matrix multiplication emerges by combining this temporal summation over 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}5 with the two-dimensional spatial indexing of the grid (Kim et al., 2024).

This optical MAC is therefore not a reprogrammable weight multiplication in the usual photonic sense. It is a direct interference measurement between two traveling optical signals of equal carrier frequency and controlled relative phase. The accumulation occurs physically as energy or photon counts at the detectors, and there is no re-encoding of intermediate electronic results back into optics within the array (Kim et al., 2024).

3. Processing-element structure and inverse-designed photonic modules

Each processing element comprises four functional blocks: a waveguide crossing, two fan-out branches, a freeform 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}6 beam splitter, and an out-coupling or detection stage. The crossing allows the vertical and horizontal buses to intersect with negligible crosstalk. The two branches tap a controlled fraction of the vertical and horizontal optical power for the local MAC while passing the remainder onward to downstream cells. The target power split for a branch of index 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}7 is approximately 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}8 for the tapped branch and 3.5μm×3.5μm3.5\,\mu\mathrm{m}\times 3.5\,\mu\mathrm{m}9 for the through path, with adjustments for realistic insertion loss. The beam splitter imposes equal power division and a deliberate 4×44\times40 phase relation needed for homodyne operation (Kim et al., 2024).

In the demonstrated array-scale configuration, the two outputs of the beam splitter are routed to vertical grating couplers and emitted out of plane, where a CMOS sensor or another two-dimensional imager measures the differential energy. The same design notes that a more integrated implementation could replace those free-space outputs with on-chip photodiodes. In either case, the local cell operation is identical: tap two traveling inputs, phase-align and amplitude-balance them, interfere them, then integrate the power difference (Kim et al., 2024).

A defining feature of the design is the use of adjoint-based inverse design for all submodules. The crossing, branches for indices 4×44\times41, beam splitter, and vertical grating coupler are each synthesized within a 4×44\times42 square by optimizing the silicon-versus-silica permittivity distribution to match target complex 4×44\times43-parameters. The phase targets are referenced to an effective propagation phase 4×44\times44, with the beam splitter alone carrying the intentional 4×44\times45 deviation. This compactness is the stated alternative to long directional couplers, multimode interference sections, and ad hoc tapers and bends, which are described as large and difficult to scale in dense two-dimensional grids (Kim et al., 2024).

The modules are implemented on silicon-on-insulator. The silicon slab waveguide has thickness 4×44\times46 and width 4×44\times47, with SiO4×44\times48 as buried oxide and cladding. Because each functional block is confined to a 4×44\times49 square, the full cell pitch is reported as 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}0, corresponding to a per-MAC footprint of 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}1 (Kim et al., 2024).

4. Simulated demonstration, normalization, and performance envelope

The reported system-level validation is a GPU-accelerated finite-difference time-domain simulation of a 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}2 photonic systolic array in Tidy3D. The demonstrated operation is an outer product, which corresponds to matrix multiplication with only one temporal element. The specific example uses real vectors

6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}3

encoded as Gaussian optical pulses with full width at half maximum 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}4 and frequency width 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}5 (Kim et al., 2024).

For each cell 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}6, the powers at the two detector outputs are integrated over a detection window to produce energies 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}7 and 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}8, and the differential output is defined as 6.2 PMACs/mm2/s6.2~\mathrm{PMACs}/\mathrm{mm}^2/\mathrm{s}9. Two normalization strategies are reported. A global normalization uses

Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}0

with Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}1 denoting one-hot vectors. A differential normalization divides each cell’s response by its own one-hot calibration value Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}2. In the illustrated example, the mean absolute error relative to the ground-truth outer product is Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}3 with simple normalization and Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}4 with differential normalization. Over a dataset of Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}5 random input pairs, the mean absolute error decreases from Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}6 to Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}7 under differential normalization (Kim et al., 2024).

The computation-density estimate is explicitly theoretical. With a cell pitch of Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}8, the area per MAC is approximately Y=WX\mathbf{Y}=\mathbf{W}\mathbf{X}9, giving about W\mathbf{W}0 MACs per W\mathbf{W}1. Assuming optical pulse trains can be streamed every W\mathbf{W}2 without overlap, corresponding to W\mathbf{W}3, the resulting throughput is W\mathbf{W}4 MACs per W\mathbf{W}5 per second, or W\mathbf{W}6 (Kim et al., 2024).

The same source states the conditions under which this estimate ceases to hold. It assumes sufficient optical bandwidth and low dispersion to support W\mathbf{W}7 pulses, detectors or CMOS imagers able to integrate over the required duration, and no system-level input/output bottlenecks. Dispersion analysis of a single MAC indicates that when Gaussian pulse width shrinks below roughly W\mathbf{W}8, corresponding to a frequency width approaching approximately W\mathbf{W}9, significant distortion and mismatch arise in the two detection ports. Extremely short pulses around A\mathbf{A}0 degrade accuracy, while sufficiently long optical pulses robustly perform scalar multiplication. The practical usable bandwidth is stated to be limited by the submodules’ transmission spectra, about A\mathbf{A}1 fractional bandwidth. Readout is also asymmetrical with respect to the optical carrier rate: the optical data rate could in principle be in the terahertz regime, but imaging frame rates are limited to megahertz, so the design is especially suited to long pulse trains and large matrices integrated over many pulses (Kim et al., 2024).

5. Relation to transformer attention and adjacent systolic-array research

The immediate machine-learning motivation is the attention score computation in transformers. Attention involves two dynamic matrices, A\mathbf{A}2 and A\mathbf{A}3, and requires the product A\mathbf{A}4. In weight-stationary photonic accelerators, this pattern is awkward because one multiplicand is usually embedded in static photonic parameters, forcing either reconfiguration overhead or partial offloading to electronics. In the all-optical systolic array, A\mathbf{A}5 and A\mathbf{A}6 can instead be treated symmetrically as optical streams: one mapped to the vertical channels, the other to the horizontal channels, so that the array computes A\mathbf{A}7 or an equivalent ordering depending on encoding convention. The subsequent softmax and multiplication by A\mathbf{A}8 are not implemented inside the reported optical array and are identified as functions that could be handled electronically from the detected outputs (Kim et al., 2024).

The architectural significance of this distinction becomes clearer when compared with recent electronic systolic-array work on attention. "SystolicAttention: Fusing FlashAttention within a Single Systolic Array" proposes an enhanced systolic array, FSA, that executes the entire FlashAttention forward pass inside one array by adding on-array support for rowwise maxima, exponentiation via exp2 and piecewise-linear interpolation, rowsums, and scaling. On a A\mathbf{A}9 array, that design reports average attention FLOPs/s utilization improvements of B\mathbf{B}0 over AWS NeuronCore-v2 and B\mathbf{B}1 over Google TPUv5e, at about B\mathbf{B}2 area overhead (Lin et al., 15 Jul 2025).

For all-optical systolic arrays, that electronic result does not supply an optical implementation of softmax, but it does sharpen the system-level question. A plausible implication is that optical matrix multiplication alone is insufficient for high-utilization attention acceleration if intermediate scores must repeatedly leave the array for external nonlinear and reduction stages. The optical proposal directly addresses the first half of the attention pipeline—dynamic-dynamic matrix multiplication—and the FSA result indicates how tightly coupled reductions and nonlinear transforms matter for end-to-end utilization (Lin et al., 15 Jul 2025).

6. Calibration, implementation constraints, and unresolved issues

The reported architecture is fully optical only in a specific and limited sense. The linear algebra operation itself is performed by optical interference and optical transport, but the accumulated result exists as detector charge or image-sensor energy. Electronics integrate and read out that charge; they do not perform the multiply–accumulate step. This distinction addresses a frequent ambiguity in the phrase “all-optical.” The design does not claim a fully photonic end-to-end attention engine, nor does it eliminate electronic detection, calibration, or final post-processing (Kim et al., 2024).

Calibration is central because residual amplitude nonuniformity, reflections, and imperfect submodule behavior lead to unequal responses across cells. The global normalization constant B\mathbf{B}3 provides a first-order correction, while per-cell one-hot calibration yields materially lower mean absolute error. The design therefore assumes a calibration phase in which one-hot patterns can be injected to measure each cell’s gain and offset, followed by digital correction or pre-distortion of input amplitudes. The control burden during normal operation is otherwise reduced: no active phase shifters are required in use, and the array is passive apart from optical sources and detectors (Kim et al., 2024).

Several open issues are explicit. Scaling beyond the numerically verified B\mathbf{B}4 array requires maintaining coherence, power uniformity, and calibration under fabrication tolerances and thermal fluctuations. Larger arrays also intensify insertion-loss management along long waveguides and increase the difficulty of keeping branch-induced power distribution uniform. Fast, low-noise detector integration compatible with the target density remains a practical challenge, whether by CMOS imaging or by more integrated photodiodes such as graphene or InGaN/GaN nanowires. The present formulation is also restricted to real-valued multiplication; extending to complex, signed, or mixed-precision arithmetic in optical-friendly formats remains open, as does embedding nonlinear functions such as softmax or activation functions in or near the photonic domain (Kim et al., 2024).

These constraints define the current position of all-optical systolic arrays in the broader photonic-computing landscape. The demonstrated principle is not simply high-throughput optical transport, but a specific co-design of data encoding, freeform photonic routing, coherent interference, and calibrated differential detection. Within that scope, the reported system establishes that a weight-free, output-stationary photonic systolic array for matrix–matrix multiplication is feasible; beyond that scope, its main unresolved questions concern scaling, readout, and integration into complete machine-learning pipelines (Kim et al., 2024).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to All-Optical Systolic Arrays.