---
title: 'VR-PTOLEMAIC: VR Spatial Audio Evaluation'
url: https://www.emergentmind.com/topics/vr-ptolemaic
type: topic
---

# VR-PTOLEMAIC: VR Spatial Audio Evaluation

VR-PTOLEMAIC is a virtual reality evaluation system for perceptual testing of spatial audio algorithms. It integrates a visually rendered virtual seminar room with head-tracked binaural playback of audio that has been convolved in real time with Ambisonic room impulse responses, and it implements the MUSHRA evaluation methodology inside VR. Assessors can teleport among 25 simulated listening positions in a virtually recreated seminar room, rotate their head naturally under 6DoF tracking, and compare simulated acoustic responses from sound field reconstruction algorithms against actually recorded second-order Ambisonic room responses, all convolved with various source signals. The system was evaluated through an extensive testing campaign, and the reported results indicate that the VR platform effectively supports the assessment of spatial audio algorithms, with generally positive feedback on user experience and immersivity [2508.00501].

## 1. System definition and scope

VR-PTOLEMAIC was designed to bridge traditional subjective evaluation methods, such as laboratory-based MUSHRA tests, with an immersive, interactive VR environment that more closely matches how listeners experience spatial audio in situ. Its core objective is not merely to reproduce a binaural listening test inside a headset, but to place the assessor in a visually reconstructed room where seat-dependent room responses, head motion, and stimulus comparison are all part of a single evaluation loop. In place of a single sweet spot or a purely binaural HRTF-based test without visual context, the framework allows assessors to move among 25 pre-defined listening positions distributed in the room and to compare reconstructed fields against measured room responses at each position [2508.00501].

The methodological centerpiece is a MUSHRA-like design with hidden references and anchors. This places VR-PTOLEMAIC within established perceptual-evaluation practice while altering the presentation layer and interaction model. A common misconception is that the system substitutes visual immersion for acoustic reference. In fact, the explicit and hidden references are created from measured second-order Ambisonic room impulse responses from HOMULA-RIR, and the simulated conditions are judged relative to those measured responses. This suggests that the framework aims to tighten the ecological validity of the test without discarding conventional reliability controls [2508.00501].

A further design choice is architectural decoupling. Visual rendering runs on an all-in-one headset, whereas audio processing runs in Max on a PC, with communication over OSC/UDP. This deployment model avoids the requirement for a VR-ready PC for graphics while preserving real-time interaction between the tracked listener and the audio engine [2508.00501].

## 2. Virtual environment and interaction model

The virtual environment is a visually reconstructed 3D model of the seminar room from the HOMULA-RIR dataset. Two spheres behind the main desk mark the sound source positions used during the measurement campaign. The 25 listening positions correspond to chair locations in the room, and movement is constrained to teleportation among these predefined spots. Head rotation is fully tracked, so the auditory scene follows the listener’s orientation even though translational motion is discretized by seat selection [2508.00501].

The Unity camera continuously sends 3D positional and rotational coordinates via OSC to the audio engine at frame rate. Seat identifiers are used to select the appropriate Ambisonic room impulse response set for the current position. Exact seat coordinates are not listed, and the manuscript does not reproduce the room’s numeric reverberation-time or clarity-index values, although it notes that HOMULA-RIR reports such properties. Visual and auditory scenes are aligned by seat-selection messages that load the corresponding A-RIR set for the selected position [2508.00501].

The interface is organized around two principal interaction surfaces. One is a seat-selection screen with buttons for the 25 positions. The other is a MUSHRA-style evaluation GUI with sliders and stimulus tabs, along with source selection, quick access to seat selection, and an information button providing attribute descriptions. The GUI is attached to the controller to support convenient positioning and activation during the session. The hardware platform comprises a Meta Quest 2 HMD with 6DoF inside-out tracking, Sennheiser HD380 Pro closed-ear headphones, and a Behringer HA400 headphone amplifier. Because playback is over closed headphones, the test room does not require special acoustic treatment [2508.00501].

## 3. Spatial audio representation and rendering pipeline

The framework operates on measured and reconstructed second-order Ambisonic room impulse responses. In the reported setup, second-order Ambisonics uses $N=2$, so the representation has $(N+1)^2=9$ Ambisonic channels. The acoustic pressure field in spherical coordinates is written as

$$
p(r,\theta,\phi,k) = \sum_{n=0}^{N} \sum_{m=-n}^{n} a_{nm}(k)\, j_n(kr)\, Y_n^m(\theta,\phi),
$$

where $j_n$ are spherical Bessel functions, $Y_n^m$ are spherical harmonics, and $a_{nm}(k)$ are the Ambisonic coefficients. In VR-PTOLEMAIC, the reference material consists of measured second-order A-RIRs from HOMULA-RIR, whereas the test material includes reconstructed A-RIRs produced by the spatial audio algorithms under evaluation [2508.00501].

The source material comprises anechoic mono excerpts of approximately $10$–$15$ s from EARS and URMP. For each selected source signal $x(t)$ and A-RIR $h(t)$, convolution is performed with HISSTools multiconvolve~ according to

$$
y(t) = x(t) * h(t) = \int_{-\infty}^{\infty} x(\tau)\, h(t-\tau)\, d\tau.
$$

Rotations and binaural decoding are then carried out in real time with SPARTA ambiBIN. Generic binaural rendering can be summarized as

$$
y_{L/R}(t) = \sum_{c=1}^{C} x_c(t) * hrtf_{c,L/R}(t),
$$

where $c$ indexes Ambisonic or virtual loudspeaker channels and $hrtf_{c,L/R}$ are head-related impulse responses for the left and right ears. SPARTA performs the appropriate HOA rotation and decoding to generate the binaural outputs [2508.00501].

The overall pipeline is therefore seat-conditioned, source-conditioned, and head-orientation-dependent. Measured and reconstructed A-RIRs are imported into Max buffers, the current seat determines which A-RIR set is loaded, and the selected stimulus is convolved and then binaurally decoded before playback over headphones. When the anchor condition is required, a low-pass filter with cutoff frequency $f_c = 3.5$ kHz is applied to the reference. This preserves much of the spatial structure while deliberately degrading timbral fidelity [2508.00501].

## 4. MUSHRA implementation and stimulus design

VR-PTOLEMAIC implements a five-stimulus evaluation layout for each attribute. The explicit reference and hidden reference are identical measured conditions; the remaining stimuli include two sound field reconstruction outputs and one low-quality anchor.

| Stimulus | Role | Construction |
|---|---|---|
| Explicit reference | Ground-truth baseline | Test audio convolved with measured second-order A-RIRs from HOMULA-RIR |
| Hidden reference $S_{\text{ref}}$ | Reliability check | Same as explicit reference but unlabeled |
| Non-parametric $S_A$ | Reconstructed condition | Amplitude matching method attributed to Abe et al. (2022) |
| Parametric $S_1$ | Reconstructed condition | Virtual miking parametric methodology attributed to Pezzoli et al. (2020) |
| Low audio quality $S_{lp}$ | Anchor | Low-pass filtered reference with cutoff frequency $3.5$ kHz |

The four rated attributes were Basic Audio Quality, Localizability, Spatial Quality, and Timbral Quality. Basic Audio Quality used a $0$–$100$ slider in accordance with ITU-R BS.1534-3 MUSHRA guidance. The other attributes were rated on clearly labeled ordinal scales: Localizability ranged from “More difficult” to “Easier,” and Spatial Quality and Timbral Quality each ranged from “Low Quality” to “High Quality.” Familiarization preceded assessment, with assessors learning the VR controls, teleportation, and GUI and listening to sample stimuli. The assessment itself comprised three trials based on different excerpts: female speech, male speech, and an instrumental track. At each trial and seat, assessors could switch among the five stimuli and rate the four attributes [2508.00501].

The manuscript gives a scoring formalization while emphasizing that the reported analysis is descriptive. A per-seat, per-trial MUSHRA score may be written as

$$
M_{i,s,a}(p,q)=r_{i,s,a}(p,q),
$$

where $r$ is the slider value or mapped ordinal score. A pooled score over seats and trials is

$$
M_{i,s,a} = \frac{1}{PQ}\sum_{q=1}^{Q}\sum_{p=1}^{P} r_{i,s,a}(p,q).
$$

Group means and confidence intervals can then be computed across assessors, although the manuscript itself reports aggregated boxplots and does not detail formal inferential tests such as ANOVA or Tukey HSD. It also does not explicitly report randomized ordering of stimuli within the GUI; instead, stimulus access is user-controlled through interface tabs, which differs from the usual randomized presentation emphasis in traditional MUSHRA practice [2508.00501].

## 5. Experimental study, behavioural logging, and reported findings

The user study involved 15 assessors, of whom 13 were male and 2 female, with average age $26.9$ years and standard deviation $4.0$. All but one held a university degree, all had prior musical experience, three were new to VR, and no hearing impairments were reported. Following ITU-R MUSHRA practices, four assessors were excluded because they rated the hidden reference below a threshold for more than 15% of test items, although the exact threshold value was not specified in the manuscript [2508.00501].

The data-collection layer included behavioural logging through OSC. Logged variables included time tracking, selected seat, evaluated attribute, stimulus changes, teleportation frequency, and time spent per seat. This logging makes the system not only an auditory rating platform but also a tool for analyzing how assessors explore the room during evaluation. Teleportation visualizations based on circle size for time spent and grayscale for teleport frequency indicated meaningful navigation patterns. Participants tended to spend more time in frontal positions, likely offering clearer spatial cues for localization and scene assessment, while still exploring other seats [2508.00501].

The reported findings are qualitative and descriptive rather than inferential. Participants could consistently differentiate among reconstruction methods across attributes, which indicates that the system surfaces perceptual differences between the non-parametric amplitude matching condition, the parametric virtual miking condition, the measured reference, and the anchor. Most participants found the interface intuitive and reported that VR helped them perceive differences in spatial attributes. Some reported mild discomfort due to headset weight during prolonged sessions. The paper reports aggregated boxplots per attribute but does not provide mean MUSHRA scores or variances explicitly in the text, and it does not report confidence intervals, effect sizes, Cronbach’s alpha, or inter-rater agreement [2508.00501].

## 6. Limitations, methodological caveats, and significance

Several limitations are explicit. First, the system is limited to second-order Ambisonic content, which implies nine-channel A-RIRs. The manuscript notes that higher-order content could improve spatial resolution and timbral stability at higher frequencies. Second, binaural decoding uses generic HRTFs rather than individualized ones, so listener-specific mismatch may increase variance in localization and externalization. Third, the OSC/UDP integration and externalized audio processing may introduce latency, but the manuscript does not quantify latency figures. Tight latency control is identified as important for VR realism [2508.00501].

A further caveat concerns ecological validity. Although the use of measured RIRs anchors the test, headphone playback is not the same as loudspeaker playback in the actual room. Likewise, the room navigation model is interactive but discretized: assessors can teleport among pre-defined seats rather than move continuously through the space. This suggests that the framework should be understood as an immersive evaluation surrogate rather than a literal duplication of in-room listening conditions. Methodologically, the reliance on descriptive statistics and the absence of explicit inferential analysis leave the reported outcomes primarily at the level of comparative usability and perceptual discriminability rather than formal hypothesis testing [2508.00501].

Within those boundaries, the system’s significance lies in combining measured second-order A-RIR references, multi-position assessment, controller-based MUSHRA interaction, head-tracked binaural rendering, and behavioural logging in a single research platform. The stated future directions are to enhance the UI and tracking analytics, integrate richer behavioural metrics, release the simulation framework as an open research tool, and extend the platform to higher-order Ambisonics, personalized HRTFs, and more diverse sound field reconstruction algorithms and settings. A plausible implication is that VR-PTOLEMAIC occupies a methodological niche between conventional lab-based perceptual audio testing and fully interactive immersive evaluation, with particular value in studies where spatial perception depends on seat position, head orientation, and exploration strategy [2508.00501].

Source: https://www.emergentmind.com/topics/vr-ptolemaic