---
title: Spatially Selective Active Noise Control
url: https://www.emergentmind.com/topics/spatially-selective-active-noise-control-ssanc
type: topic
---

# Spatially Selective Active Noise Control

Searching arXiv for recent papers on Spatially Selective Active Noise Control and closely related formulations.
Spatially Selective Active Noise Control (SSANC) denotes a family of active noise control formulations in which the controller is designed to satisfy an explicitly spatial objective rather than merely minimizing residual pressure at one microphone or an unweighted sum over several microphones. In the literature summarized here, spatial selectivity appears in several forms: suppression over a continuous target region \(\Omega\), preservation of desired sound from selected directions at a listener’s eardrum, prioritization of “critical” points within a controlled volume, and directional filter selection in reverberant environments. Across these variants, the common principle is to embed spatial structure directly into the optimization problem through kernels, linear constraints, target-response constraints, or direction-conditioned filter banks [2202.04807, 2208.09997, 2507.05657].

## 1. Scope and defining characteristics

Conventional ANC is commonly formulated to achieve maximal sound reduction regardless of the incident direction of the sound, or to minimize the mean-square error \(E\{e^2(n)\}\) at one or more error microphones. The SSANC literature identifies several limitations of this paradigm. In volumetric ANC, traditional multi-point schemes suppress average acoustic energy within a volume by minimizing the sum of squared error signals at all control microphones, but uniform weighting treats all locations equally and offers no way to prioritize “critical” points such as a listener’s ears over “auxiliary” regions [2507.05657]. In personal and wearable ANC, state-of-the-art schemes often cancel both desired sound and noise and then reconstruct the desired sound, which can introduce latency and spectral or binaural-cue distortions and can increase control effort [2208.09997]. In spatial ANC over a continuous region, pointwise control at a sparse set of microphones does not by itself guarantee attenuation throughout the interior of the target region \(\Omega\) [2202.04807, 2303.16021].

A common misconception is that SSANC refers to a single architecture. The literature instead uses the term for multiple design strategies. In continuous-region formulations, SSANC seeks to reduce the regional acoustic potential energy
\[
J=\int_{\Omega}|u(r)|^2\,dr,
\]
where \(u(r)=p(r)+s(r)\) is the total field formed by primary and secondary sound fields [2202.04807]. In open-fitting hearables, SSANC seeks to attenuate undesired noise at the eardrum while preserving desired speech arriving from a selected direction, typically through a constraint on the desired-source transfer path rather than by post hoc reconstruction [2401.07681, 2507.12122]. In volumetric ANC, spatial selectivity is obtained by solving a linearly constrained minimum variance problem that enforces exact responses at selected locations while minimizing residual variance elsewhere [2507.05657]. In learning-based directional SSANC, selectivity is realized by estimating the direction of arrival (DoA) and switching among pre-trained fixed filters matched to discrete directions [2601.06981].

## 2. Optimization frameworks and canonical constraints

The principal SSANC formulations can be organized by how spatial information enters the cost function. The table summarizes representative forms drawn directly from the cited literature.

| Formulation | Canonical expression | Spatial role |
|---|---|---|
| Continuous-region SSANC | \(J=\int_{\Omega}|u(r)|^2dr\) | Suppress noise over a target region |
| Hard-constrained hearable SSANC | \(H^T u=f\) or \(H(\mathbf{q}+\mathbf{G}\mathbf{w})=\bm{\updelta}_{\Delta}\) | Preserve desired-source transfer exactly |
| Soft-constrained hearable SSANC | \(J_{\text{time}}(\mathbf{w})=E\{e^2(n)\}+\beta\|\mathbf{w}\|_2^2+\mu\|H[\mathbf{q}+\mathbf{G}\mathbf{w}]-\delta_{\Delta}\|_2^2\) | Trade off speech distortion and noise reduction |
| Volumetric LCMV ANC | \(\min \mathbf{w}^T R_x \mathbf{w}\;\text{s.t.}\; C^T\mathbf{w}=d\) | Prioritize critical points within a volume |

In the volumetric linearly constrained minimum variance formulation, the filtered reference covariance is \(R_x=E\{x(n)x^T(n)\}\), the constraint matrix \(C\) collects steering vectors for the selected constraint points, and the target responses are encoded in \(d\). The closed-form solution is
\[
w_{\mathrm{opt}}=R_x^{-1}C(C^T R_x^{-1}C)^{-1}d.
\]
This expression separates exact constraint satisfaction from residual-variance minimization: \((C^T R_x^{-1}C)^{-1}d\) enforces the desired gains at the chosen points, while the \(R_x^{-1}\) factor shapes the remaining degrees of freedom to minimize spillover noise elsewhere in the volume [2507.05657].

In hearable SSANC, a central distinction is between hard and soft constraints. Hard-constrained designs impose exact preservation of a desired speech component, for example by requiring
\[
\mathbf{H}(\mathbf{q}+\mathbf{G}\mathbf{w})=\bm{\updelta}_{\Delta},
\]
where \(\mathbf{H}\) is built from relative impulse responses (ReIRs), \(\mathbf{q}\) encodes the leakage path, and \(\bm{\updelta}_{\Delta}\) selects the desired delay [2505.10372]. Soft-constrained designs relax this exact equality into a quadratic penalty controlled by a frequency-independent parameter \(\mu\), so that conventional ANC and hard-constrained SSANC become limiting cases:
\[
\lim_{\mu\to 0} w_{\text{soft}}(\omega)=w_{\text{ANC}}(\omega),\qquad
\lim_{\mu\to \infty} w_{\text{soft}}(\omega)=w_{\text{hard}}(\omega).
\]
This formulation makes the noise-reduction versus speech-preservation trade-off explicit [2507.12122].

Adaptive implementations follow the same logic. Xiao, Xu and Zhao derived a Frost-style projected LMS recursion,
\[
w(n+1)=P[\,w(n)-\mu G^T x(n)e(n)\,]+q,
\]
where the projection matrix \(P\) and offset \(q\) enforce the spatial constraint after each unconstrained ANC update [2208.09997]. In volumetric LCMV ANC, a constrained FxLMS update is obtained via a generalized sidelobe canceller decomposition or stochastic-gradient treatment of instantaneous Lagrange multipliers, with the update projected into the nullspace of the constraint matrix so that adaptation does not violate the chosen spatial constraints [2507.05657].

## 3. Continuous-region and volumetric SSANC

A major branch of SSANC addresses suppression over a continuous spatial region rather than at a few discrete microphones. In the formulation of spatial ANC based on individual kernel interpolation, the primary field \(p(r)\), the secondary field \(s(r)\), and the total field \(u(r)=p(r)+s(r)\) are defined on \(\Omega\subset\mathbb{R}^3\). Measured error-microphone signals are used together with a positive-definite kernel \(\kappa(\cdot,\cdot)\) to interpolate the field throughout the region:
\[
\hat u(r)=z(r)^T e,\qquad z(r)=[(K+\lambda I)^{-1}]^T\kappa(r),
\]
which converts the regional cost into a quadratic form \(J=e^H A e\) [2202.04807].

Directional weighting is introduced through
\[
\kappa(r_1,r_2)=\frac{1}{4\pi}\int_{S_2}\gamma(\xi)e^{j k \xi^T(r_1-r_2)}d\xi,\qquad
\gamma(\xi)=\exp[\beta\,\eta^T\xi],\ \beta\ge 0,
\]
so that prior information on the primary-source direction \(\eta\) can bias the interpolation toward plane-wave components arriving from that direction. Larger \(\beta\) concentrates the kernel more strongly along \(\eta\) [2202.04807]. The same paper identifies an important limitation of “total-kernel-interpolation”: applying the same directional weighting to the total field can be suboptimal because the secondary sources generally lie in directions different from the primary source. The proposed remedy is “individual kernel interpolation,” in which the primary and secondary fields are estimated separately:
\[
\hat d=e-\hat G y,\qquad s=\hat G y,
\]
with separate kernels tailored to the source direction of each component. The resulting quadratic cost in \(\hat d\) and \(y\) yields a normalized least-mean-square update for the multichannel control filter [2202.04807].

Arikawa, Koyama and Saruwatari proposed a related but sensor-economical approach in which the primary field inside \(\Omega\) is interpolated from reference microphones outside the target region rather than from many interior error microphones. Kernel-ridge regression gives
\[
\hat p(r,n)=z(r)^T x_n,\qquad z(r)=(K+\lambda I_R)^{-1}\kappa(r),
\]
while the secondary field is modeled by free-field Green’s functions. Integrating the estimated total-field energy over \(\Omega\) produces offline-computable matrices \(A_{yy}\), \(A_{yx}\), and \(A_{xx}\), from which the fixed spatial filter follows as
\[
W=-A_{yy}^{-1}A_{yx}.
\]
To compensate interpolation error, they further introduced a hybrid cost
\[
C(n)=\gamma^n J(n)+\|e_n\|_2^2,
\]
which transitions from the interpolated-field objective to conventional multichannel ANC using only a small number of interior error microphones [2303.16021].

A further extension adds explicit control of exterior radiation. In kernel-interpolation-based spatial ANC with exterior radiation suppression, the interior acoustic potential energy \(J_{\rm int}(w)\) is supplemented either by a penalty term \(\beta J_{\rm ext}(w)\) or by an inequality constraint \(J_{\rm ext}(w)\le\gamma\), where
\[
J_{\rm ext}(w)=w^H A_{\rm ext} w
\]
represents the radiated exterior power. This yields two adaptive algorithms: an exterior-penalized NLMS and a projected NLMS that enforces the radiation bound [2303.16389].

The reported numerical results establish the significance of these formulations. In the individual-kernel study, the proposed method achieved \(\sim 20\) dB reduction at 200 Hz versus \(\sim 16\)–\(17\) dB for total-kernel interpolation, outperformed the comparison methods across 100–600 Hz, and remained robust under \(\pm 6^\circ\) azimuth and \(\pm 3^\circ\) elevation perturbations of the primary direction [2202.04807]. In the reference-interpolation study, a fixed kernel-interpolation filter achieved \(\sim 14\) dB uniformly over \(\Omega\) from the first iteration at 400 Hz, and the hybrid “NLMS w/ Fixed-KIR” rose to \(\sim 18\) dB after 10 000 iterations, outperforming conventional NLMS by 5–10 dB across 100–500 Hz [2303.16021]. In the exterior-suppression study, both proposed methods maintained exterior power at 50% of the vanilla NLMS while trading a slight interior loss of less than 2 dB across 100–1000 Hz [2303.16389]. Volumetric LCMV ANC complements these continuous-region approaches by allowing exact prioritization of selected points while retaining broadband regional suppression [2507.05657].

## 4. Hearable SSANC and distortionless preservation

In hearable systems, SSANC is formulated around a listener-specific control point, typically an inner microphone near the eardrum. The canonical architecture comprises \(K\) outer microphones, one inner error microphone, one loudspeaker, and a controller that drives anti-noise through the secondary path. A representative time-domain model writes
\[
e(n)=p(n)+(\mathbf{G}\mathbf{w})^T\mathbf{x}(n)=(\mathbf{q}+\mathbf{G}\mathbf{w})^T\mathbf{x}(n),
\]
where \(\mathbf{x}(n)\) stacks the outer-microphone signals and a leakage estimate, \(\mathbf{G}\) is the secondary-path convolution matrix, and \(\mathbf{q}\) selects the leakage term [2505.10372, 2605.17407].

The key conceptual move in hearable SSANC is that the desired sound is preserved physically rather than reconstructed. In the multi-channel augmented-eyeglasses system of Xiao, Xu and Zhao, the disturbance at the error microphone is \(d(n)=s(n)+v(n)\), where \(s(n)\) is the desired source and \(v(n)\) is the undesired sound. The constraint
\[
H^T u=f,\qquad u=G w+\tilde\delta,
\]
holds the transfer from the desired direction to the error microphone identical to the uncontrolled response. The minimization of \(E\{e^2(n)\}\) is therefore restricted to the nullspace of \(H^T\), so that the algorithm cancels only disturbance components orthogonal to the desired path [2208.09997]. The paper explicitly relates this mechanism to an ANC-augmented minimum-power distortionless response beamformer.

Target-signal definition and delay are critical design variables. For open-fitting hearables, one can define the desired target either as the desired component at the error microphone,
\[
t_{\rm err}(n)=p_s(n),
\]
or as a delayed desired component at a reference microphone,
\[
t_{\rm ref}(n)=x_{{\rm ref},s}(n-\tau_{\rm ref}).
\]
When the reference-microphone target is imposed at the error microphone, causality requires a delay equal to the acoustic arrival difference between the reference and error microphones:
\[
\tau_{\rm opt}=d_{\rm acoustic}/c,\qquad
\Delta_{\rm opt}=\lceil f_s d_{\rm acoustic}/c\rceil.
\]
The simulations show that the error-microphone target achieves optimal performance without delay, whereas the reference-microphone target is infeasible at zero delay and performs best at the acoustic delay [2401.07681].

Acausal optimization modifies this picture by allowing anti-causal taps in the ReIRs. In the acausal hearable formulation, the outer-microphone signals depend on relative impulse responses indexed over \(l\in[-L_a,L_h-1]\), and the constraint
\[
\mathbf{H}(\mathbf{q}+\mathbf{G}\mathbf{w})=\bm{\updelta}_{\Delta}
\]
is solved by a regularized least-squares system. When \(L_a=0\), the formulation reduces to the purely causal SSANC solution of prior work; increasing \(L_a\) introduces up to \(L_a\)-sample anti-causal taps. The reported simulations show substantially lower speech distortion without sacrificing noise reduction. In a scenario with speech at \(0^\circ\) and an interferer at \(45^\circ\), the causal design with \(L_a=0\) yielded \(\mathrm{SD}_{\mathrm{intellig}}\approx -2\) dB, \(\mathrm{NR}\approx 21\) dB, and \(\mathrm{SNR}\approx +5\) dB, whereas the acausal design with \(L_a=22\) yielded \(\mathrm{SD}_{\mathrm{intellig}}\approx -28\) dB, \(\mathrm{NR}\approx 21\) dB, and \(\mathrm{SNR}\approx +17\) dB [2505.10372].

The empirical record of hearable SSANC is correspondingly strong. On six-channel augmented eyeglasses, the proposed hard-constrained system improved residual SNR from \(-13.9\) dB a priori to \(+15.2\) dB, achieved noise reduction of \(29.1\) dB, and attained a speech-distortion index of \(-25.1\) dB above 100 Hz, while using only 2% of the secondary-source energy required by one reference method and substantially less than the decoupled alternative [2208.09997]. In open-fitting hearables, the error-microphone target with \(\Delta=0\) achieved approximately \(16.1\) dB noise reduction and the reference-microphone target achieved approximately \(15.4\) dB at \(\Delta\approx 4\) samples, which matched the acoustic delay [2401.07681]. These results support the view that spatial selectivity in hearables is primarily a problem of imposing the correct distortionless spatial constraint under strict causality and secondary-path limitations.

## 5. Directional, fixed-filter, and learning-based SSANC

A different realization of SSANC replaces online coefficient adaptation with direction-conditioned filter selection. In directional selective fixed-filter ANC, a multi-reference feedforward ANC setup uses \(J\) spatially distributed microphones, one secondary loudspeaker, and one error microphone. Rather than adapting \(w_j(n)\) at run time, the method pre-trains a library \(\{w_j^{(k)}\}_{k=1}^K\) of fixed filters and seeks, ideally, to choose
\[
\hat k=\arg\min_k e_k^2(n),
\]
which is approximated by selecting the filter whose pre-training direction best matches the live DoA [2601.06981].

The DoA estimation stage uses a convolutional neural network operating on short-time Fourier transform features from the reference array. For each frame, the tensor
\[
R\in\mathbb{R}^{J\times F\times T\times 2}
\]
collects magnitude and phase across channels, frequencies, and time frames. The network maps this tensor to azimuth and elevation class probabilities,
\[
(\hat p_{\rm azim},\hat p_{\rm elev})=\mathrm{CNN}(R;\Theta^*),
\]
followed by class selection via \(\arg\max\). The architecture comprises three convolutional modules with \(3\times 3\) kernels and \(\{16,32,64\}\) feature maps, each followed by group normalization, ReLU, and \(2\times 2\) max-pooling, then an adaptive average-pool and two parallel fully connected softmax heads. The network was trained by a multi-task cross-entropy loss on 0.5 s noise excerpts convolved with room impulse responses spanning three room sizes, nine \(RT_{60}\) values from 0.1 to 0.9 s, eight array positions per room, SNR uniformly in \([30,50]\) dB, and both synthetic and UrbanSound8K noises [2601.06981].

The fixed-filter library itself encodes the spatial selectivity. For each discrete direction \((\theta_i,\phi_k)\) on an \(A\times B\) angular grid, classical FxLMS is run offline with broadband training noise from that direction to yield the steady-state filters \(\{w_j^{(i,k)}\}\). At runtime, once the CNN outputs \((\hat a,\hat b)\), the controller simply switches to
\[
w_j(n)\leftarrow w_j^{(\hat a,\hat b)},\qquad j=1,\dots,J.
\]
Because the system avoids per-sample adaptation, the runtime complexity reduces to \(J\cdot L\) multiplies plus the CNN inference every \(T_f\) ms [2601.06981].

The reported performance establishes that this architecture is viable in reverberant environments, not only in free field. The CNN achieved approximately 96.4% azimuth accuracy and approximately 91.0% elevation accuracy at SNR\(\ge 30\) dB, with 0.03 M parameters, 119.9 M MACs, and a per-frame runtime of 7.8 ms on CPU. In a simulated tetrahedral 4-microphone configuration, directional SFANC achieved about 15 dB broadband reduction after 0.5 s for a source at \((120^\circ,30^\circ)\), compared with about 11 dB for conventional FxLMS, about 12 dB for standard SFANC, and about 10 dB for GFANC; when the source moved to \((0^\circ,-30^\circ)\), directional SFANC stabilized around 13 dB within 0.3 s while SFANC and GFANC amplified the residual by about 2 dB; with real washing-machine noise from \((110^\circ,-15^\circ)\), it outperformed all baselines by about 4–6 dB and adapted gracefully to an out-of-library direction [2601.06981]. This suggests that direction-conditioned fixed-filter selection can serve as a computationally efficient form of SSANC when rapid response is more important than continuous online adaptation.

## 6. Design trade-offs, robustness, and practical limitations

The SSANC literature repeatedly emphasizes that spatial selectivity is obtained by spending degrees of freedom strategically, which makes design trade-offs unavoidable. In volumetric LCMV ANC, the practical recommendation is to use only the smallest necessary set of \(K\) critical points so that \(M-K\) degrees of freedom remain for global noise minimization; increasing \(K\) tightens spatial selectivity but reduces residual-variance minimization capability [2507.05657]. In continuous-region kernel methods, dense boundary microphone placement and directional kernels improve regional estimation, but larger \(\Omega\) or higher frequencies require more reference microphones or richer kernels to capture the spatial complexity [2202.04807, 2303.16021]. In exterior-radiation suppression, the penalty formulation converges faster, whereas the constrained formulation guarantees an a priori radiation bound [2303.16389].

For hearables, the most explicit trade-off is between speech distortion and noise reduction. The soft-constrained design introduces a frequency-independent parameter \(\mu\) in
\[
J_{\text{time}}(\mathbf{w})=E\{e^2(n)\}+\beta\|\mathbf{w}\|_2^2+\mu\|H[\mathbf{q}+\mathbf{G}\mathbf{w}]-\delta_{\Delta}\|_2^2,
\]
and the simulations demonstrate a broad plateau of useful operating points. With \(\mu=0\), conventional ANC achieved about 21.6 dB noise reduction but maximum speech distortion, \(\Delta\mathrm{SNR}\) of only about 4.4 dB, PESQ loss of about \(-0.12\), and ESTOI gain of \(+0.06\). In the hard-constrained limit \(\mu\to\infty\), speech preservation became nearly perfect, with \(\mathrm{SD}\approx -31.5\) dB, but noise reduction decreased to about 12.1 dB. An intermediate setting such as \(\log_{10}\mu=-2\) yielded \(\mathrm{NR}\approx 20.7\) dB, \(\mathrm{SD}\approx -14.7\) dB, \(\Delta\mathrm{SNR}\approx 17.2\) dB, PESQ gain of \(+0.54\), and ESTOI gain of \(+0.39\), substantially outperforming the hard-constrained design in \(\Delta\mathrm{SNR}\), PESQ, and ESTOI [2507.12122].

Robustness to modeling error, especially secondary-path mismatch, is another central concern. Robust soft-constrained SSANC for hearables addresses this by averaging the cost over a library of \(M\) candidate secondary-path estimates \(\{\widehat{\mathbf{G}}_i\}_{i=1}^M\). The average-cost design yields the closed-form solution
\[
\mathbf{w}^*
=-
\Bigl[
\overline{\bm{\Phi}_{rr}}
+\mu \frac{1}{M}\sum_{i=1}^M \widehat{\mathbf{G}}_i^T \mathbf{H}^T\mathbf{H}\widehat{\mathbf{G}}_i
\Bigr]^{-1}
\Bigl[
\overline{\bm{\phi}}
-\mu \frac{1}{M}\sum_{i=1}^M \widehat{\mathbf{G}}_i^T\mathbf{H}^T(\alpha\bm{\delta}_{\Delta}-\mathbf{H}\mathbf{q})
\Bigr].
\]
The study reports that the matched case delivered the best mean noise reduction, about 20–25 dB, with a 5–95 percentile spread within \(\pm 0.5\) dB; the mismatched case could exhibit a spread exceeding 5–6 dB; and the robust average-cost design stayed within 1–2 dB of the matched-case mean while shrinking the 5–95 percentile range back toward \(\pm 1\) dB. Real-time validation on a dSPACE SCALEXIO platform reproduced the simulated speech and noise spectra at the eardrum [2605.17407].

Computationally, SSANC ranges from modest to demanding depending on the formulation. Volumetric LCMV ANC can precompute \(P_{\rm null}\) and \((C^T C)^{-1}\) when \(C\) is time-invariant, leaving real-time cost of approximately \(O(ML+MK^2)\) per sample [2507.05657]. Acausal hearable optimization requires forming and inverting matrices of size \((K+1)L_w\), with a stated cost of \(\mathcal{O}(((K+1)L_w)^3)\) per update, though the block-diagonal structure of \(\mathbf{G}\) can be exploited [2505.10372]. Directional fixed-filter SSANC reduces runtime adaptation cost to fixed convolution plus intermittent CNN inference [2601.06981]. These differences underscore that “spatial selectivity” does not imply a single complexity profile; the implementation burden depends on whether the spatial constraint is enforced continuously by adaptation, approximately by soft penalties, or discretely by filter selection.

Several limitations remain explicit in the literature. Very dense reverberant fields can blur DoA cues, and the directional fixed-filter method does not yet model source-to-array distance variability [2601.06981]. Reference microphones placed far from \(\Omega\) observe only low-order field components, increasing interpolation error at higher frequencies or in the presence of scattering [2303.16021]. Ill-conditioning of \(A_{yy}\), \(C^T C\), or related matrices necessitates regularization such as \(\lambda I\), \(\epsilon I\), or diagonal loading [2202.04807, 2507.05657]. These are not peripheral implementation details; they are structural consequences of asking an ANC system to shape an acoustic field selectively rather than merely reduce error power wherever it is measured.

Taken together, the literature portrays SSANC as a unifying perspective on ANC design. Whether formulated through kernel interpolation over \(\Omega\), distortionless constraints for desired directions, LCMV constraints at critical points, or DoA-conditioned filter banks, the essential aim is the same: minimize unwanted acoustic energy subject to spatial requirements that reflect the priorities of the application.

Source: https://www.emergentmind.com/topics/spatially-selective-active-noise-control-ssanc