---
title: Compressive Single-Pixel Imaging
url: https://www.emergentmind.com/topics/compressive-single-pixel-imaging-spi
type: topic
---

# Compressive Single-Pixel Imaging

Searching arXiv for recent and foundational papers on compressive single-pixel imaging to support the article.
{"query":"single-pixel imaging compressive sensing complementary modulation Hadamard Morlet wavelet Fourier arXiv", "max_results": 10}
{"query":"all:(single-pixel imaging) AND all:(compressive sensing)", "max_results": 10}
Compressive single-pixel imaging (SPI) is an indirect imaging paradigm in which a scene is encoded by a sequence of known spatial patterns and measured by a bucket detector rather than by a focal-plane array. In its standard formulation, the unknown image \(x\in\mathbb R^{N}\) is probed by a measurement matrix \(\Phi\in\mathbb R^{M\times N}\), producing measurements \(y=\Phi x+n\), with \(M\ll N\) in the compressive regime and recovery relying on sparsity or compressibility in a transform domain \(\Psi\) [2601.15248]. Across the literature, this framework is instantiated with digital micromirror devices (DMDs), spatial light modulators (SLMs), complementary or differential acquisition, deterministic or randomized bases, and increasingly with learned optical encoders and learned decoders, enabling operation in visible, infrared, terahertz, hyperspectral, and multiphoton settings [2203.09802].

## 1. Forward model and compressive formulation

The common mathematical core of compressive SPI is the linear forward model
\[
y = A x + n,
\]
or equivalently \(y=\Phi x+n\), where \(x\) is the vectorized scene, \(A\) or \(\Phi\) is the pattern matrix, \(y\) is the vector of bucket measurements, and \(n\) models additive noise [1707.03164]. For full-basis sampling with an orthogonal basis and \(M=N\), direct inversion or a pseudo-inverse recovers \(x\); for compressive SPI, \(M<N\), so the inverse problem is underdetermined and must be regularized [2601.15248].

The standard compressive-sensing formulation writes \(x=\Psi x'\) or \(x=\Psi s\), with \(x'\) or \(s\) sparse in the basis \(\Psi\), yielding
\[
y=\Phi \Psi x' + n.
\]
Recovery is then posed through \(\ell_1\)-constrained or \(\ell_1\)-penalized optimization, for example
\[
\min \|x'\|_1 \quad \text{subject to} \quad \|y-\Phi\Psi x'\|_2\le \varepsilon,
\]
or through total-variation regularization when the image gradient is sparse [2203.09802]. The tutorial treatment of SPI gives the usual compressive-sensing scaling \(M \gtrsim K\log(N/K)\) when the unknown has \(K\) significant coefficients in the transform domain [2601.15248].

This formulation extends naturally beyond monochrome 2-D imaging. In single-pixel Fourier-transform interferometry and related hyperspectral systems, the operator becomes a Kronecker product of a spatial coding basis and a spectral Fourier basis, with recovery in a joint spatial-spectral sparsifying transform [1810.13200]. In time-domain terahertz SPI, the same spatial model is applied slice-wise to time-resolved waveforms, so that each spatial coefficient carries a reconstructed temporal transient [1911.09545]. These variants preserve the same inverse-problem structure while enlarging the object from a 2-D image to a datacube or a field indexed by time or wavelength.

## 2. Pattern design and sampling ensembles

Pattern design is central because it determines measurement efficiency, optical throughput, reconstruction cost, and noise sensitivity. A basic division is between randomized ensembles, deterministic orthogonal bases, and bases tailored either to image statistics or to a particular optical transfer function.

Deterministic orthogonal coding is dominated by Hadamard-type constructions. The SPI tutorial distinguishes natural ordering, Walsh ordering, and cake-cutting ordering, and notes that \(\pm1\) Hadamard rows can be implemented on binary hardware by complementary measurements or by using the all-ones row as a reference [2601.15248]. In high-efficiency variants, the deterministic basis is reordered so that low-complexity patterns appear first. Cake-cutting Hadamard SPI sorts patterns by connected-component count and reports reconstructions of \(1024\times1024\) images at a sampling ratio of \(0.2\%\), with \(O(N\log N)\) runtime and \(O(N)\) memory via the fast Walsh-Hadamard transform [1903.11175]. Origami pattern construction similarly imposes a deterministic ordering of orthogonal \(\{\pm1\}\) patterns and reports operation down to \(0.5\%\) sampling for \(128\times128\) images [1903.11432].

Fourier-domain sampling replaces random projection by direct acquisition of informative spatial-frequency coefficients. The eSPI method projects sinusoidal patterns, exploits Hermitian symmetry of the Fourier spectrum of real images, and samples only the most informative frequency bands. In the reported \(128\times128\) simulation, \(10\%\) coverage required 1638 patterns and achieved RMSE \(=0.061\), while random patterns with linear correlation gave RMSE \(=0.215\) and random patterns with compressive sensing gave RMSE \(=0.115\); the real experiment used 1635 total patterns and obtained visually faithful reconstructions at about \(30\) dB SNR [1504.03823].

A different statistical strategy is to randomize within a structured feature family. Czajkowski et al. proposed Morlet-wavelet-correlated random patterns obtained by convolving white Gaussian noise with 2-D Morlet wavelets, so that measurements probe localized joint spatial-frequency features. In numerical tests on \(512\times512\) images at \(4\%\) compression, average PSNR was \(25\) dB for Morlet-random patterns, compared with \(22\) dB for noiselets and \(14\) dB for Walsh-Hadamard; in optical experiments on \(256\times256\) images at \(6\%\) compression, TV reconstruction yielded PSNR \(27\) dB for binarized Morlet-random patterns, versus \(25.5\) dB for noiselets and \(18\) dB for Walsh-Hadamard [1709.07739].

At the opposite extreme, high-resolution sparse-scene SPI uses highly compressed map-based binary patterns displayed at full DMD resolution. One framework acquires \(3002\) patterns for \(1024\times768\) images, corresponding to \(CR\approx0.38\%\), with differential, binary, non-adaptive sampling and a map construction designed to preserve information about multiple image partitions and an unknown field of view [2206.02510]. A related high-resolution system uses approximately \(3.2\times10^3\) binary patterns at a compression ratio of \(0.41\%\), where each pattern is about \(50\%\) on, and reports \(6.8\) Hz acquisition at \(22\) kHz DMD operation for sparse scenes [2509.01497].

## 3. Complementary, differential, and balanced measurements

A distinctive feature of many SPI systems is that the measurement is not taken directly as \(y_i=\langle \phi_i,x\rangle\), but as a difference between complementary or successive patterns. This is especially natural on a DMD, because each micromirror directs light into one of two reflection arms. If \(P_i^+\) and \(P_i^-\) denote the two reflected patterns, then
\[
P_i^+(c,d)+P_i^-(c,d)=1,
\]
and the first-order differential measurement
\[
y^+-y^-=(A^+-A^-)\Psi x' + (e^+-e^-)
\]
eliminates the large DC term and doubles the useful AC contrast, provided the two optical arms are perfectly balanced [2203.09802].

The practical problem is that the two arms are rarely perfectly matched. The system of “Secondary complementary balancing compressive imaging with a free-space balanced amplified photodetector” addresses gain and bias mismatch by acquiring two complementary differential measurements per pattern pair, \(B_1\) and \(B_2\), and then forming the second-order difference \(B_1-B_2\). In the reported derivation, all bias terms and common-mode noise cancel, leaving a balanced differential measurement proportional to \(\hat A \Psi x'\) up to a scalar factor \(k'\) [2203.09802]. The implementation uses a silicon free-space balanced amplified photodetector with \(5\) mm active diameter, directly outputting the photocurrent difference between the two DMD arms. Under \(25\%\) sampling in simulation, the method matches the ideal dual-arm case and significantly outperforms an unbalanced dual-arm system even for \(1\%\)–\(5\%\) gain error; experimentally, at \(40\%\) sampling, its MSSIM equals conventional compressive sensing under full sampling, and DC offset and common-mode noise are reduced to within the digitizer’s noise floor [2203.09802].

Single-arm complementary modulation predates this dual-arm balancing strategy and is especially important in microscopy. In complementary microscopic ghost imaging, each random binary pattern is followed by its logical complement, and the effective bucket signal is
\[
y_k=\tfrac12\,(y'_{\text{odd},k}-y'_{\text{even},k}),
\]
with an effective sensing row taking values in \(\{+0.5,-0.5\}\). This subtraction removes constant background, including lamp drift and dark counts, and halves the variance of i.i.d. noise. At \(50\%\) sampling, the reported PSNR is approximately \(32\) dB for complementary compressive sensing, versus approximately \(17\) dB for conventional compressive sensing without subtraction [1501.06002].

Differential acquisition is also the basis of differential ghost imaging and of photon-starved Hadamard SPI. The unified comparison of SPI algorithms presents the DGI estimator
\[
\hat x = \langle y_i a_i\rangle - [\langle y\rangle/\langle s\rangle]\langle s_i a_i\rangle,
\]
and notes that it is used experimentally for better SNR [1707.03164]. In single-photon cake-cutting Hadamard SPI, each \(\pm1\) pattern is implemented by displaying a binary pattern and its complement in immediate succession; the differential count \(y^+-y^-\) restores the desired signed inner product and suppresses common-mode noise and drift [1903.11175].

## 4. Reconstruction algorithms

Reconstruction methods in compressive SPI range from non-iterative correlation and direct inverse transforms to convex optimization, regularized generalized inverses, unfolded optimization networks, and end-to-end learned decoders. The algorithmic comparison paper places these approaches in a single framework and emphasizes a three-way trade-off between capture efficiency, computational complexity, and robustness to noise [1707.03164].

Correlation-based methods are the simplest. Traditional ghost imaging and differential ghost imaging reconstruct by averaging pattern–measurement products, with DGI providing better SNR in experiments [1707.03164]. When the sensing matrix has orthonormal rows or can be orthogonalized, direct inversion becomes feasible. In Morlet-pattern SPI, a pre/post-processing SVD converts a non-orthogonal matrix into an orthonormalized form suitable for standard CS solvers, while the precomputed pseudoinverse
\[
\hat X = M^+ Y
\]
is reported to be about \(50\times\) faster than CS-TV and suitable for real-time reconstruction at moderate SNR [1709.07739]. In terahertz SPI, Zanotto et al. used direct Hadamard decoding \(\hat x = H_{\text{sub}}^T y\) for rapid reconstruction when only the first \(M\) ordered Hadamard rows are measured [1911.09545]. In eSPI, the recovered Fourier coefficients are simply filled into a sparse spectrum and transformed by one inverse FFT [1504.03823].

Iterative model-based methods remain the dominant reference algorithms. The unified comparison includes alternating projection, conjugate gradient descent for the normal equations, Poisson maximum likelihood, and TV-regularized compressive sensing implemented with ADMM-like updates [1707.03164]. The reported conclusion is precise: for comparable reconstruction accuracy, TV requires the least measurements and the least running time for small-scale reconstruction; CGD and AP run fastest in large-scale cases; and TV and AP are the most robust to measurement noise [1707.03164]. TVAL3, NESTA, SPGL1, and other basis-pursuit-denoising solvers recur throughout the literature, especially when the sensing basis is not exactly orthogonal or when edge preservation is critical [1709.07739]. For high-resolution sparse-scene imaging, differential Fourier-domain regularized inversion (D-FDRI) replaces iterative \(\ell_1\)-minimization by a precomputed generalized inverse in the Fourier domain, followed by a map-based refinement that zeros sectors classified as empty and rescales non-empty sectors [2206.02510].

Deep learning enters SPI in two distinct ways: learned decoders with fixed measurement patterns, and end-to-end learning of both patterns and reconstruction. In block compressive sensing, the BCS-UNet architecture applies a fixed Bernoulli block matrix \(A\) to each \(B\times B\) block and then reconstructs images of arbitrary size above a minimum block size through an UpsampleNet plus UNet pipeline; the reported model is trained on natural-image datasets yet reconstructs SPI-acquired binary transmissive targets with PSNR/SSIM within \(1\) dB/\(0.02\) of simulation [2207.06746]. In deep-learned orthogonal-basis SPI, a modified deep convolutional autoencoder treats the encoder weight matrix as the measurement basis and imposes binary and orthogonality regularizers; at \(6.25\%\) compression on \(64\times64\) images, inference takes about \(3\) ms per frame, and non-binary learned bases achieve mean SSIM around \(0.81\) at additive white Gaussian noise level \(5\times10^{-4}\), above classical TV and Fourier baselines [2205.08736].

More recent systems combine a fast initial inverse with a refinement stage. One \(1024\times768\) framework first computes \( \hat x_0=\Phi^+ y \) and then applies either a few ISTA/FISTA iterations or a U-net-inspired enhancement network. At \(CR=0.41\%\) and \(6.8\) Hz acquisition, the initial image yields PSNR \(17\)–\(22\) dB and SSIM \(0.50\)–\(0.70\); iterative enhancement yields \(28\)–\(32\) dB and \(0.85\)–\(0.95\); neural enhancement yields \(30\)–\(35\) dB and \(0.90\)–\(0.98\) [2509.01497]. Self-supervised SPI reconstruction pushes the optimization prior back into the network architecture. SISTA-Net unfolds ISTA into a fidelity module and a proximal mapping module, combines CNN and visual state-space modeling, and trains only from the measurements and the physical forward model. The reported gains are about \(2.6\) dB in simulation and \(3.4\) dB average PSNR in far-field underwater experiments [2603.29732].

## 5. Hardware realizations and modality-specific extensions

The optical front-end of compressive SPI is unusually flexible because the detector is non-pixelated and the spatial coding can be implemented in many different ways. The canonical visible-light bench uses a DMD or SLM to display binary or grayscale patterns, collection optics to image the modulated field onto a photodiode or PMT, and synchronized acquisition electronics [2601.15248]. Within that basic architecture, the literature spans microscopy, multiphoton imaging, terahertz time-domain spectroscopy, hyperspectral interferometry, stochastic analog modulators, and passive diffractive encoders.

Microscopy is an early application because single-pixel detection is compatible with low light and with wavelengths where detector arrays are expensive. In complementary microscopic ghost imaging, the microscopic image is focused onto a \(768\times1024\)-mirror DMD rather than projecting speckles onto a tiny object, and only the \(+12^\circ\) reflection arm is collected by a Hamamatsu PMT. The system reconstructs large images row by row and is reported to be robust under ultra-weak halogen illumination, with the unused DMD arm left open for future infrared sampling [1501.06002]. In wide-field multiphoton microscopy with single-pixel detection, Wijesinghe et al. combined temporal focusing, an SLM, and Morlet-basis compressive sensing in the TRAFIX configuration. The reported faithful reconstructions reach \(M/N\approx0.125\), with root-mean-square error below \(15\%\) up to seven scattering lengths and with an eight-fold reduction in pattern number translating into an eight-fold faster acquisition and up to eight-fold lower cumulative light dose [1907.02272].

Terahertz and hyperspectral systems broaden the dimensionality of the forward model. In time-domain terahertz SPI, a DMD encodes binary masks onto an \(800\) nm beam, which photoexcites a silicon wafer and modulates THz transmission with approximately \(95\%\) depth; Hadamard-coded measurements reconstruct spatial amplitude, per-pixel temporal waveforms, time-of-flight thickness maps, and spectral images, with recognizable images down to about \(40\%\) measurements [1911.09545]. Moshtaghpour, Bioucas-Dias, and Jacques, and related SP-FTI work, integrate Hadamard spatial coding with Fourier-transform interferometry along the optical-path-difference axis. Their formulations use \(\Phi_{\rm had}\otimes\Phi_{\rm dft}\), assume sparsity in 3-D Haar or wavelet domains, and adopt variable-density or multilevel sampling to reduce both measurement rate and light exposure relative to Nyquist FTI [1810.13200]. Numerical results for SP-FTI report reconstruction SNR around \(20\) dB at \(10\%\) of full measurements under measurement SNR \(20\) dB, while the broader hyperspectral FTI study reports SRE about \(18.3\) dB with multilevel-guided sampling at measurement-use ratio approximately \(0.11\), compared with approximately \(0.6\) dB for uniform-density sampling at the same rate [1809.00950].

SPI can also be realized without a programmable DMD. The stochastic spatial light modulator replaces the electronic modulator by a vibrating transparent chamber filled with opaque particles; an overhead camera thresholds the random particle arrangement into a binary mask, and the resulting Bernoulli-like sensing matrix is used with TVAL3 reconstruction. The paper emphasizes application domains in which no analog to optical SLMs exists, including THz, X-ray, ultrasound, and non-optical lensless imaging [1810.08694]. At the opposite technological end, Wang et al. recast compressive SPI as a learned linear intensity transformation \(y=Hx+n\), pre-train the target matrix jointly with a shallow decoder ANN, and then optimize a wavelength-multiplexed spatially incoherent diffractive optical processor to approximate that matrix. The resulting encoder is passive and static, with acquisition limited by LED switching and detector readout rather than by SLM refresh [2603.21456].

A further shift is from reconstructing hyperspectral datacubes to performing spectral tasks directly from compressed measurements. HyPIS uses two spectral encoders with wavelength-dependent sine and cosine transfer functions and a DMD for spatial-temporal illumination. Three synchronized single-pixel detectors capture the unencoded, cosine-encoded, and sine-encoded signals, from which two phasor images \(G(x)\) and \(S(x)\) are reconstructed. The reported system generates \(64\times64\) phasor frames at \(2.68\) fps, achieves \(100\%\) classification accuracy over six object classes, and remains accurate under low light and strongly non-uniform illumination [2604.01801].

## 6. Performance regimes, trade-offs, and recurrent misconceptions

The performance envelope of compressive SPI is defined by a persistent coupling between measurement count, optical SNR, modulator speed, and reconstruction cost. The algorithm-comparison study makes this explicit: TV has the best capture efficiency at small scale, AP and CGD are faster at large scale, and TV and AP are most robust to Gaussian measurement noise [1707.03164]. This already rules out a common simplification that “best” SPI reconstruction can be specified independently of image size, pattern family, and noise model.

A second recurrent misconception is that SPI is intrinsically restricted to very low spatial resolution. Historically, many papers reported \(32\times32\) to \(256\times256\) reconstructions, but this is not a hard limit of the modality. Full-resolution DMD imaging at \(1024\times768\) has been demonstrated with acquisition times on the order of a fraction of a second when the scene is sparse or the field of view is limited but a priori unknown [2206.02510]. A separate framework reports \(1024\times768\) acquisition at \(6.8\) Hz with a compression ratio of \(0.41\%\), but explicitly states that only spatially sparse scenes admit high-fidelity recovery at that extreme compression, while dense scenes lose fine structure [2509.01497]. The appropriate conclusion is not that resolution is no longer a problem, but that the limiting variable has shifted from detector pixel count to scene structure, compression ratio, and compute budget.

A third misconception is that compressive SPI is synonymous with random projection plus generic \(\ell_1\)-minimization. The literature now includes importance-sampled Fourier coefficients [1504.03823], deterministic reordered Hadamard bases such as cake-cutting and origami [1903.11175], Morlet feature-matched random ensembles [1709.07739], deep-learned orthogonal or binary bases [2205.08736], and fixed passive diffractive encoders learned jointly with a shallow decoder [2603.21456]. This suggests that “compressive SPI” is better understood as a design space in which the sensing operator is co-optimized with the optical hardware, the expected image class, and the intended reconstruction algorithm.

Noise handling is similarly more nuanced than a simple preference for differential acquisition. Complementary subtraction can remove DC terms, background drift, dark counts, and common-mode noise, but only if the differential architecture itself is stable. The dual-arm balanced-photodetector work shows that imbalance as small as \(1\%\)–\(5\%\) gain error can noticeably degrade uncorrected dual-arm performance, while secondary complementary balancing largely restores the ideal case; it also notes that extremely high imbalance, above \(100\%\) gain error, may reduce SNR if residual noise dominates [2203.09802]. Differential measurement is therefore not merely a hardware convenience but a calibration-sensitive part of the inverse model.

Documented development paths are correspondingly diverse. Explicitly proposed directions include faster spatial modulators or LED arrays with complementary coding for high-speed operation, extension to IR and THz bands where balanced detectors are available but pixel arrays are not, hybrid optical-digital co-design of static diffractive encoders, and spectral-task pipelines that bypass high-resolution hyperspectral cubes entirely [2203.09802]. The field therefore evolves not toward a single canonical SPI architecture, but toward specialized compressive imagers whose sensing basis, detector physics, and reconstruction prior are matched to modality, photon budget, and downstream task.

Source: https://www.emergentmind.com/topics/compressive-single-pixel-imaging-spi