FieldCore: Multiplexed Photonic Tensor Core
- FieldCore is a fully multiplexed photonic tensor core that exploits five native multiplexing dimensions for massive parallelism and efficient tensor operations.
- It leverages inverse-designed silicon photonics and a programmable optical crossbar to achieve uniform, high-speed computation across up to 1,800 parallel input streams.
- The architecture is validated through diverse AI and signal processing tasks, achieving 69.12 TOPS while maintaining high precision and scalability.
FieldCore is a fully multiplexed photonic tensor core architecture that exploits all five native multiplexing dimensions—wavelength, radio-frequency, guided mode, time, and space—within a single optical field to achieve massive computational parallelism and high aggregate throughput for tensor operations. Leveraging inverse-designed silicon photonics, FieldCore enables uniform, programmable computation across all multiplexed channels, with validated high-speed performance on a range of computational and AI inference workloads. It accommodates up to 1,800 parallel input streams and demonstrates peak aggregate compute throughput of 69.12 tera operations per second (TOPS), establishing a scalable platform for fully multiplexed photonic tensor computing (Sun et al., 24 Apr 2026).
1. Native Multiplexing Dimensions and Architecture
FieldCore leverages the full suite of photonic multiplexing dimensions to maximize compute density and parallelism:
- Time (): Serial symbol streams up to 120 GBaud support fine-grained data parallelism.
- Radio-Frequency Subcarriers (): Up to 100 RF channels per optical carrier, yielding vectorized data within each wavelength.
- Wavelength Channels (): Dense wavelength-division multiplexing (WDM) across the C-band (≥40 nm), organized into banks.
- Guided Modes (): Multiple spatial modes—demonstrated with TE₀ and TE₁, extensible to higher-order modes—enable parallel streams within a fiber.
- Space (input-output ports, ): spatial input and output ports configured as a crossbar.
The architecture constructs tensors hierarchically across these axes:
- Scalar regime: Single mode, wavelength, subcarrier, port, and time stream.
- Vector regime: Adding RF subcarriers, yielding -length vectors.
- 2D Tensor: Inclusion of WDM produces 0 slices per mode/port.
- 3D Tensor: Mode-division multiplexing yields 1 arrays per port.
- 5D Tensor: Utilizing all input and output ports realizes 2.
The block-level schematic sequences input modulation (MZM for RF up-conversion and WDM), polarization control, mode MUX, programmable multiplier-adder crossbar (using dual-port and single-port MZIs with designed splitting ratios and uniform analog weights), and demultiplexing stages (mode deMUX, WDM demux, photodetection, and electrical RF demux).
2. Inverse-Designed Photonic Building Blocks and Tensor Computation
FieldCore's central computational operation is a mode-1 tensor–matrix multiplication, with the input tensor
3
and a programmed weight matrix
4
The element-wise parallel computation is given by
5
or compactly, 6.
Technological implementation utilizes:
- Mode MUX/deMUX: Inverse-designed for broadband operation (< -20 dB crosstalk), 120 nm² pixel resolution, and <1 dB insertion loss.
- Digital-metamaterial splitters/crossings: Provide compact, low-crosstalk signal routing.
- Tunable couplers/adders (MZIs): Facilitate programmable fan-out and accumulation in the crossbar with designed splitting ratios.
- Multipliers: Single-port MZIs plus mode-insensitive phase shifters deliver uniform weighting with <1% mismatch between TE₀/TE₁ and >30 dB extinction ratio.
- Photodetectors: Supporting up to 100 GHz bandwidth, followed by electrical RF demultiplexing.
The number of independent parallel input streams is given by 7, with demonstrated scalability to 1,800 parallel streams.
3. Compute Throughput, Performance Metrics, and Precision
Compute throughput is defined by parallel MACs across all multiplexed channels. The aggregate per-port baudrate is 8, and the compute throughput is
9
Expressed in TOPS units,
0
In a representative throughput-maximized configuration (1, 2, 3, 4, 5, 6 GBaud), FieldCore achieves approximately 69.12 TOPS.
Benchmarks validate:
- Arithmetic precision: Mode invariance within 2.9% for 7 vs. 8 weighted states. MAC precision remains above 5 bits up to 120 GBaud; error distributions remain Gaussian with 9. Wavelength uniformity yields <0.7 bits variation from 1530–1565 nm. Addition precision with RF scaling (0, 1 GBaud) remains ∼5 bits.
- Parallel stream capacity: Up to 1,800 parallel input streams are projected by maximizing RF, WDM, and MDM axes.
4. Experimental Demonstrations and Application Benchmarks
FieldCore's versatility is established through a spectrum of AI and signal processing tasks:
- Arithmetic Operations: Uniform computation validated across channels and multiplexing axes, including high-speed symbol rates (up to 120 GBaud).
- Image Convolution: Grayscale edge convolution with 2×2 kernels across four directions achieves mode-averaged PSNRs: 30.5 dB at 10 GBaud, 23.4 dB at 70 GBaud; intensity scaling over 2 yields PSNR of 30.11 dB (3) and 28.54 dB (4). RGB convolution using three FDM subcarriers provides PSNR >27 dB per channel.
- Handwritten Digit Recognition (MNIST): 200 fully multiplexed inputs (5) processed in parallel with single-optical-layer convolution achieve 93.8% average accuracy versus a 95.0% digital baseline.
- Hyperspectral Classification (Indian Pines, 200 bands): Channel-parallel mapping of 2×2 windows and 200 bands processes up to 200 parallel streams, with four shared 2×2 kernels and 12-class digital classification, yielding 90.8% overall accuracy and robust precision down to 4 bits (FieldCore equivalent ~5.16 bits).
- Mechanical Fault Diagnosis: Monitoring hundreds of vibration signals (Case Western Reserve dataset), FieldCore processes 200 parallel channels with four shared length-4 kernels, producing 95.1% accuracy and stable precision above 5 bits (FieldCore equivalent ∼4.77 bits).
5. Scaling Laws, Trade-Offs, and Design Considerations
Multiplicative scaling in FieldCore is governed by the product of all available channelization axes: 6. Key design considerations include:
- Bandwidth Limitations: RF parallelism is currently limited by detector and oscilloscope bandwidth but can scale to 200 GHz+ with integrated electronics.
- Mode Scaling: Scaling to higher-order modes and leveraging polarization offer further expansion, contingent on managing crosstalk and fabrication variance via inverse design.
- WDM and Source Scaling: Use of microcomb sources and wavelength-insensitive couplers can extend 7 to >50.
- Spatial Fabric Growth: Demonstrated on a 4×4 crossbar, FieldCore is compatible with foundry scaling to larger 8 fabrics.
- Precision–Throughput Trade-off: Increasing baudrate reduces dynamic range; however, error-aware training techniques enable robust inference and classification with 4–6 bit precision without significant performance loss.
6. Summary and Outlook
FieldCore realizes a general, scalable photonic tensor-core by combining all five native multiplexing axes of guided light within an integrated silicon-photonic platform. Its inverse-designed elements ensure mode-insensitive, broadband, high-fidelity tensor computation and weighted summation across hundreds to thousands of co-propagating channels. FieldCore's architecture supports ultra-high-speed MAC computation (95 bit precision, 120 GBaud), multi-dimensional convolution, and parallel inference workloads, validated by application-level demonstrations in hyperspectral imaging, mechanical fault diagnosis, and digit classification. The system supports up to 1,800 independent input streams and achieves projected throughput of 69.12 TOPS, substantiating the paradigm of fully multiplexed photonic tensor computing (Sun et al., 24 Apr 2026).