Papers
Topics
Authors
Recent
Search
2000 character limit reached

Ambisonics Upscaling: Methods and Advances

Updated 14 July 2026
  • Ambisonics Upscaling is a technique for enhancing spatial fidelity by transforming low-order Ambisonics signals into higher-order representations using spherical harmonics.
  • It employs diverse methods such as classical source-field directional emphasis, beamforming from sparse arrays, and data-driven neural super-resolution.
  • Key insights include addressing underdetermined inverse problems, managing hardware constraints, and balancing computational cost with perceptual improvements.

Searching arXiv for papers on Ambisonics Upscaling and closely related methods. Ambisonics Upscaling (AU) denotes methods that increase the order of an Ambisonics representation, or otherwise improve the effective spatial fidelity of a lower-order or constrained Ambisonics signal, without requiring a conventional higher-order capture setup. In the strict sense used by the directional-emphasis literature, AU converts a low-degree Ambisonics representation into a higher-degree representation by operating on a source-field description on the 2-sphere (Kleijn, 2018). In more recent work, the term also covers beamforming-based order up-scaling from sparse smartphone arrays (You et al., 3 Jun 2026), FOA-to-HOA3 neural super-resolution (Nawfal et al., 1 Aug 2025), and cascaded diffusion-based FOA-to-3rd-order generation (Milstein et al., 30 Sep 2025). Closely related lines of work use reduced representations, residual channels, or covariance-based transcoding to synthesize or reconstruct spatial detail at the decoder or renderer, although several of those papers explicitly distinguish themselves from classical order-increasing AU (Hold et al., 2024, Politis et al., 16 Jun 2026).

1. Formal problem and representational basis

Ambisonics represents a sound field in spherical harmonics. For order NN, the number of channels in 3D is (N+1)2(N+1)^2, so FOA has 4 channels and HOA3 has 16 channels (Nawfal et al., 1 Aug 2025, Milstein et al., 30 Sep 2025). A standard internal sound-field expansion is

p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),

with Bnm(k)B_n^m(k) the Ambisonics coefficients, jnj_n the spherical Bessel function, and YnmY_n^m the spherical harmonics (Kleijn, 2018). AU becomes nontrivial because recovering higher-order coefficients from lower-order coefficients is underdetermined. DiffAU states this explicitly as

aN(k)=FaN(k),N>N,a_N(k) = F a_{N'}(k), \qquad N' > N,

where FF extracts the first (N+1)2(N+1)^2 channels from the higher-order signal; since FF is wide, there are infinitely many higher-order signals compatible with the FOA input (Milstein et al., 30 Sep 2025).

This inverse character is central across the literature. In physics-based AU, the missing directional detail is introduced by a structured operator defined on a source-field distribution. In model-based capture upscaling, the missing detail is inferred from steering functions, array manifolds, or residual subspaces. In data-driven AU, the missing detail is sampled or regressed from learned priors over higher-order Ambisonics. A recurrent implication is that AU is not a single algorithmic family but a class of reconstructions constrained by spherical-harmonic structure, capture geometry, or learned data distributions.

2. Classical AU as source-field directional emphasis

The canonical AU formulation in "Directional emphasis in ambisonics" treats enhancement as multiplication of a source field on the sphere by an angular weighting function (Kleijn, 2018). The source field is written as

(N+1)2(N+1)^20

with the asymptotic relation

(N+1)2(N+1)^21

Defining (N+1)2(N+1)^22, the emphasized source field becomes

(N+1)2(N+1)^23

where (N+1)2(N+1)^24 is a real, ideally nonnegative angular function (Kleijn, 2018).

The essential upscaling step follows from the fact that products of spherical harmonics can be re-expanded through a sparse matrix of Clebsch–Gordan coefficients: (N+1)2(N+1)^25 This yields the core AU equation

(N+1)2(N+1)^26

which maps a low-degree Ambisonics vector (N+1)2(N+1)^27 to a higher-degree vector (N+1)2(N+1)^28 after directional weighting (Kleijn, 2018). The paper further shows that for a fixed emphasis operator the mapping can be implemented as a single matrix multiply from (N+1)2(N+1)^29 input channels to p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),0 output channels, with runtime cost p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),1 multiplies per output sample. For a degree-1 Ambisonics input and a degree-2 emphasis operator, only four multiplies per output channel are needed (Kleijn, 2018).

The method supports both static and adaptive arrangements. Static emphasis uses a time-invariant p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),2 to emphasize a fixed direction or set of directions. Adaptive emphasis estimates a time-varying emphasis function from the signal itself, with one proposed form

p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),3

The paper also notes a projection step for idempotent adaptive rendering: because of spherical-harmonic orthogonality, the low-degree coefficients can simply be overwritten with the original ones, adding no extra computational cost (Kleijn, 2018). This establishes a precise historical meaning of AU: not merely order conversion, but directional sharpening via source-field-domain weighting and controlled re-expansion.

3. Capture-side upscaling from sparse and irregular arrays

A major contemporary use of AU arises when capture hardware cannot support conventional HOA encoding. SHB-AE addresses the smartphone case directly. Classical HOA capture assumes a well-sampled spherical microphone array with at least p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),4 microphones to encode order p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),5, whereas a typical smartphone may have only four microphones placed asymmetrically on the handset body (You et al., 3 Jun 2026). For a smartphone array with p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),6, direct HOA encoding becomes underdetermined beyond first order. SHB-AE reinterprets Ambisonics encoding as a beamformer-design problem: p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),7 and chooses one beamformer per spherical harmonic basis function so that

p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),8

Using a DSHT matrix p(r,θ,ϕ,k)=n=0m=nnBnm(k)jn(kr)Ynm(θ,ϕ),p(r,\theta,\phi,k)=\sum_{n=0}^{\infty}\sum_{m=-n}^{n} B_n^m(k)\,j_n(kr)\,Y_n^m(\theta,\phi),9, the beamformer is constrained by

Bnm(k)B_n^m(k)0

so that the designed beamformer behaves like a unit selector for the target basis component in the spherical-harmonic domain (You et al., 3 Jun 2026).

The distinctive claim of SHB-AE is that the usual bottleneck of “not enough microphones” is replaced by a bottleneck in the number of steering directions used to characterize the array manifold. The method measures or simulates steering vectors for many directions on a sphere, rather than treating microphones as the only sampling points. To reduce aliasing and stabilize high-frequency design, the paper introduces a frequency division strategy in which, above a threshold Bnm(k)B_n^m(k)1, the array manifold is replaced by its magnitude Bnm(k)B_n^m(k)2 to suppress erroneous inter-microphone phase differences (You et al., 3 Jun 2026). In the reported setup, 50 source directions are chosen using a Gaussian distribution over the sphere, with five elevation angles Bnm(k)B_n^m(k)3 and azimuth samples every Bnm(k)B_n^m(k)4, and the method successfully encodes and up-scales Ambisonics up to the fourth order with just four irregularly arranged microphones (You et al., 3 Jun 2026).

The same general problem appears in arbitrary-array work that does not present itself as classical AU but remains directly relevant to it. "Ambisonics Encoding For Arbitrary Microphone Arrays Incorporating Residual Channels For Binaural Reproduction" formulates a signal-independent linear encoder

Bnm(k)B_n^m(k)5

derives a Tikhonov-regularized optimal filter, and introduces a null-space criterion

Bnm(k)B_n^m(k)6

to predict which Ambisonics channels are accurately encodable from steering functions alone (Gayer et al., 2024). Its residual channels,

Bnm(k)B_n^m(k)7

are not standard Ambisonics channels, but they capture higher-order spatial content omitted by a truncated Ambisonics vector and substantially reduce binaural NMSE in the reported simulations (Gayer et al., 2024). This expands the capture-side understanding of AU: the target need not always be a conventional higher-order channel set; it can also be a hybrid base-plus-residual representation.

4. Data-driven FOA-to-HOA super-resolution

Neural AU makes the underdetermined inverse problem explicit and replaces analytic priors with learned conditional mappings. "Ambisonics Super-Resolution Using A Waveform-Domain Neural Network" takes a 4-channel FOA signal encoded with SN3D normalization and ACN channel ordering and outputs a 16-channel HOA3 signal in the same convention (Nawfal et al., 1 Aug 2025). The model is a modified Conv-TasNet, with FOA input, HOA3 output, 384 encoder channels, single repetition Bnm(k)B_n^m(k)8, 256 channels in the separator/upscaler stage, and total model size 1,428,764 parameters. Training uses internal datasets totaling 2000 hours of speech, music, and noise, with augmentation over source count up to 12, random azimuth and elevation, fixed distance at 1 m, 4-second excerpts, gain in the range Bnm(k)B_n^m(k)9 dB to jnj_n0 dB, and simulated rooms (Nawfal et al., 1 Aug 2025).

The evaluation protocol uses a 16,382-point spherical 180-design grid, 250 ms pink-noise sources, and positional mean squared error in dB after conventional decoding. Reported average positional errors are about 16 dB for FOA rendering, about 4 dB for native HOA3 rendering, and about 4.6 dB for the proposed FOA→HOA3 renderer (Nawfal et al., 1 Aug 2025). A multiple-stimuli listening test with a hidden reference found a significant effect of Renderer, jnj_n1, and post-hoc comparisons showed that HOA3 and the proposed renderer did not differ significantly; the median qualitative rating indicated an 80% improvement in perceived quality over the traditional rendering approach (Nawfal et al., 1 Aug 2025). The paper explicitly frames the result as a learned approximation rather than a physically exact reconstruction.

DiffAU pushes the same FOA→HOA3 objective into conditional diffusion. It decomposes the task into order-by-order steps,

jnj_n2

using two NCSN++-style score networks, one for jnj_n3 and one for jnj_n4 (Milstein et al., 30 Sep 2025). Each stage predicts only the newly required channels, consistent with

jnj_n5

so the first block predicts 5 channels and the second predicts 7 channels, for 12 predicted channels total (Milstein et al., 30 Sep 2025). The method operates in the STFT domain with a nonlinear amplitude transform

jnj_n6

and uses predictor-corrector sampling with SNR parameter jnj_n7, 30 predictor steps per diffusion block, and 1 corrector step per predictor iteration (Milstein et al., 30 Sep 2025).

On a synthetic WSJ0-based anechoic dataset with 1 to 4 speakers and disjoint speakers across splits, DiffAU is evaluated with STFT-SDR on HOA channels 5–16. Reported overall performance is jnj_n8 dB for DiffAU versus jnj_n9 dB for the PWD-CS baseline, with DiffAU ahead for every speaker count (Milstein et al., 30 Sep 2025). In a MUSHRA-style test with 9 participants, DiffAU was reported to be perceptually indistinguishable from the 3rd-order reference, while the FOA anchor averaged 28.7 (Milstein et al., 30 Sep 2025). The neural and diffusion papers therefore establish AU as a conditional generative problem rather than only a beamforming or harmonic-analysis problem.

5. Broader and adjacent formulations

Several neighboring lines of work use AU-like logic without always performing literal order increase.

Formulation Mechanism Relation to AU
"Methodology for 3D sound synthesis of directional acoustic sources by higher-order ambisonics" (Thorner et al., 2024) Multimodal horn model, dual-layer sphere of 100 virtual sensors, Lebedev projection, 3D 5th-order HOA playback with 56 loudspeakers Modeled directional source is projected into HOA and rendered
"Dynamic Real-Time Ambisonics Order Adaptation for Immersive Networked Music Performances" (Ostan et al., 1 Aug 2025) Throughput monitoring, threshold-based order switching, instantaneous or cross-faded adaptation Real-time downscaling and recovery of order under bandwidth constraints
"Perceptually-motivated Spatial Audio Codec for Higher-Order Ambisonics Compression" (Hold et al., 2024) Downmix to transport channels, perceptual metadata, exact low-order recovery, higher-order synthesis Decoder-side reconstruction from reduced representation, not classical AU
"Generalised Transcoding Framework for Arbitrary Spatial Audio Capture and Playback Formats" (Politis et al., 16 Jun 2026) Covariance fitting of primary sources plus ambience, optimal mixing matrices, residual ambience injection Generalised AU-compatible transcoding rather than simple SH order conversion
"Array-Aware Ambisonics and HRTF Encoding for Binaural Reproduction With Wearable Arrays" (Gayer et al., 15 Jul 2025) Array-aware ASM encoding plus AA-MagLS HRTF preprocessing AU-like end-to-end enhancement, not true order increase
"DynFOA" (Luo et al., 6 Feb 2026) 360° video, 3D Gaussian Splatting, conditional diffusion FOA generator Spatial-audio generation of FOA, adjacent to AU rather than higher-order upscaling

The directional-source synthesis paper is important because it shows an “upscaled” pipeline from a physically modeled horn to a reproducible HOA field. The radiated pressure is sampled on a dual-layer sphere of 100 virtual sensors distributed according to a Lebedev grid, projected into normalized spherical harmonics, converted to HOA coefficients through Williams’ method, and reproduced on a 3D 5th-order HOA spatialization sphere with 56 loudspeakers (Thorner et al., 2024). The work does not start from low-order Ambisonics, but it demonstrates AU-like source-to-field-to-loudspeaker synthesis for strongly directional sources.

Dynamic order adaptation addresses a different constraint: network throughput rather than capture geometry. The proposed controller monitors available bandwidth using YnmY_n^m0, uses the selection rule

YnmY_n^m1

and lowers or restores Ambisonics order in real time (Ostan et al., 1 Aug 2025). Under the tested 5% packet-loss condition, instantaneous order reduction preserved both BAQ and localizability better than slow cross-fade, and the paper concludes that instantaneous order reduction is preferable to a slow cross-fade under that condition (Ostan et al., 1 Aug 2025). This is not AU in the order-increasing sense, but it is part of the same spatial-resolution trade space.

The HOA codec based on HO-DirAC is even more explicit about the distinction. It does not take a low-order HOA signal and deterministically recover a higher-order representation from it. Instead, it downmixes full HOA into fewer transport channels, recovers low orders using a truncated pseudo-inverse, and synthesizes remaining higher-order components using scene parameters such as direction of arrival, diffuseness, and sector energy (Hold et al., 2024). Likewise, the nCOMPASS framework generalises AU into covariance-based transcoding: it estimates time-frequency-dependent metadata for primary source components and an ambience component, constructs target playback covariances, and derives optimal mixing matrices, with explicit support for independent rotations of capture and playback setups (Politis et al., 16 Jun 2026). The paper states that its strongest gains appear for lower-order Ambisonics and geometrically constrained microphone arrays (Politis et al., 16 Jun 2026).

The wearable-array AA-MagLS method further clarifies a common boundary condition. It improves binaural reproduction from imperfectly encoded Ambisonics by making HRTF preprocessing aware of the actual arbitrary-array encoder, but it does not create true higher-order information if the array cannot support it (Gayer et al., 15 Jul 2025). DynFOA is adjacent for the opposite reason: it generates FOA from 360-degree video using geometry- and material-aware conditional diffusion, so it is a spatial-audio generation framework rather than higher-order AU (Luo et al., 6 Feb 2026).

6. Limitations, misconceptions, and research directions

A recurrent misconception is that AU is always exact recovery of physically missing higher-order coefficients. The literature does not support that claim. DiffAU explicitly frames AU as recovery from an underdetermined inverse problem and addresses it by learning the posterior distribution YnmY_n^m2 (Milstein et al., 30 Sep 2025). The waveform-domain Conv-TasNet approach also presents FOA→HOA3 as a learned mapping whose output approaches native HOA3 quality but remains a data-driven estimate (Nawfal et al., 1 Aug 2025). Even the classical directional-emphasis operator is not coefficient recovery; it is a controlled transformation of the source-field distribution (Kleijn, 2018).

A second misconception is that higher order is always useful if it can be generated. SHB-AE reports clear reconstruction improvement up to 4th order for the tested four-microphone smartphone array, with little benefit beyond that, and the paper concludes that gains beyond 4th order were not demonstrated to be useful in that setup (You et al., 3 Jun 2026). The networked-music paper reaches a related conclusion from the transport side: under bandwidth stress, lower order can outperform a corrupted higher-order stream in both audio quality and localizability (Ostan et al., 1 Aug 2025). These results indicate that order should be treated as a resource-constrained design parameter rather than a monotonic objective.

A third misconception is that any perceptual improvement in binaural rendering constitutes AU. The codec, array-aware HRTF, and covariance-transcoding papers are careful on this point. The HO-DirAC codec is a parametric reduction-and-reconstruction system, not classical AU from low-order HOA alone (Hold et al., 2024). AA-MagLS improves the effective Ambisonics-to-binaural pipeline from wearable arrays, but the paper states that it does not create true higher-order information if the array cannot support it (Gayer et al., 15 Jul 2025). nCOMPASS can be viewed as a broad AU-compatible transcoding method, yet it reconstructs target playback covariances from estimated source-plus-ambience metadata rather than by algebraically increasing spherical-harmonic order (Politis et al., 16 Jun 2026).

Current limitations remain method-specific and substantial. SHB-AE remains limited by geometry-dependent aliasing and high-frequency manifold uncertainty (You et al., 3 Jun 2026). DiffAU reports results only in free-field, noiseless, reverberation-free conditions (Milstein et al., 30 Sep 2025). DynFOA notes approximate material estimation and experiments mostly in controlled indoor environments (Luo et al., 6 Feb 2026). The arbitrary-array residual-channel framework is formulated for binaural reproduction and states that generalizing those residuals to other rendering tasks would need additional study (Gayer et al., 2024). A plausible implication is that future AU systems will continue to hybridize three ingredients already visible across the literature: explicit array or propagation models, scene-level metadata, and generative priors.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Ambisonics Upscaling (AU).