Papers
Topics
Authors
Recent
Search
2000 character limit reached

EMORSION Protocol: Audio and Film Immersion

Updated 5 July 2026
  • EMORSION Protocol is a research framework that manipulates film audio dimensions (frequency, dynamics, and directionality) to study their impact on audience emotion and immersion.
  • The protocol employs controlled audio mix variations during cinematic screenings and gathers data using subjective surveys, physiological monitoring, and motion tracking.
  • Preliminary findings indicate that subtle changes in audio design can generate measurable differences in immersion, emotional response, and viewer synchrony.

EMORSION, short for “Examining the Impact of Audio Parameters on Emotional Responses and Immersion in Film,” is an exploratory proof-of-concept protocol for studying how film-audio design shapes audience emotion, attention, and immersion in a real cinema setting. It is structured around controlled manipulations of three audio dimensions—frequency, dynamics, and directionality—while preserving ecologically valid movie-viewing conditions through theatrical exhibition, repeated scene presentation, and a triangulated multimodal assessment framework combining self-report, physiological monitoring, and video-based motion analysis (Garcia et al., 29 May 2026).

1. Conceptual scope and research question

EMORSION is designed to address a specific question: How do targeted changes in film audio—specifically pitch/frequency, loudness/dynamics, and spatial directionality—alter viewers’ emotional responses and sense of immersion? Its stated motivation is that, although music has been studied extensively, sound effects and detailed audio-design parameters remain underexplored, especially under ecologically valid exhibition conditions with a live audience.

The protocol therefore aims to isolate the perceptual contribution of individual audio parameters without reducing the viewing experience to a small-scale laboratory task. This suggests a methodological compromise between experimental control and cinematic realism: the manipulations are narrowly specified, but the presentation context remains a functioning cinema environment rather than a simplified audiovisual test bench.

A central implication of this framing is that EMORSION is not merely a study of soundtrack preference. It is a protocol for testing whether relatively granular changes in audio design can produce measurable differences across subjective, physiological, and behavioural channels while viewers engage with narrative film scenes in conditions closer to standard exhibition practice.

2. Stimulus corpus and exhibition design

The protocol uses four film scenes spanning horror and drama/suspense, balanced between mainstream and independent productions. Scene selection followed consultation with Queen Mary School of Drama experts and two professional sound engineers. The stated criteria were a strong balance of music and sound effects including Foley and environmental sound, a stand-alone narrative, suitable emotional range for immersive viewing, and limited stylistic variability within genres (Garcia et al., 29 May 2026).

Scene Metadata Target emotion
Ford vs Ferrari (FVF) Adventure/Suspense; 2h02–2h10; 8 min Tense, wonder
A Quiet Place (AQP) Horror; 5:00–10:00; 5 min Sad, tense
I Saw the TV Glow (ISTVG) Horror; 58:45–1h04; 5 min Intrigue, tense
Decision to Leave (DTL) Suspense; 1h35–1h46; 10 min Tense, intrigue

The study reports that horror and drama/suspense were used to reduce stylistic variability while preserving genre contrast, and that one mainstream and one independent scene were selected for each genre. All scenes were between 5 and 10 minutes. Familiarity was low: only two or three participants recognized each film.

The exhibition component took place in three sessions at BLOC Studios, using a 36-speaker Dolby Atmos system and 4K projection. There were 40 participants in total: 17 male, 22 female, and 1 non-binary. Session sizes were 13, 13, and 14. In each session, participants viewed four scenes, and each scene was shown twice—once as a control mix and once as an augmented mix—for 8 presentations total per participant, comprising 4 control and 4 augmented presentations. Scene order and augmentation assignment were counterbalanced.

3. Audio manipulations and mix construction

For each scene, EMORSION created four mixes: 1 original control mix, specified as a 7.1.2 Dolby Atmos mix, and 3 augmented mixes, each manipulating only one audio axis. Across the four scenes, this yielded 16 total audio mixes. The mixes were produced in Reaper and DaVinci Resolve using factory plug-ins.

The protocol defines the manipulated dimensions as follows. Dynamics is the “manipulation of level and dynamic range via compressors, limiters, and expanders, controlling contrast between soft and loud events.” Frequency comprises “modifications to spectral and pitch-related characteristics, brightness, timbral weight, and tonal centre, using equalization, saturation, distortion, and key transposition.” Directionality is the “alteration of spatial audio distribution via stereo and 5.1 Atmos panning, affecting sound source localisation and spatialisation.”

These three dimensions correspond, respectively, to loudness contrast, pitch/spectral shape, and spatial placement/localization. Because each augmented mix changes only one axis relative to the control, the protocol is explicitly structured to probe parameter-specific effects rather than holistic soundtrack redesign. A plausible implication is that EMORSION can be used to distinguish immersive effects produced by conventional spatialization from those produced by timbral or dynamic alteration, provided the scene context is held fixed.

4. Triangulated multimodal measurement framework

EMORSION evaluates audience response through a triangulated framework combining subjective, physiological, and behavioural measures (Garcia et al., 29 May 2026).

The self-report component uses a questionnaire completed on mobile devices after each scene. The protocol description refers to a six-item questionnaire, while the analysis section describes a five-item questionnaire assessing emotional response and immersion. The measured dimensions included emotional response, emotional change over time, perceived immersion, salient element identification, and related ratings. The reported inferential procedures were ANOVA for emotional intensity ratings and salient element identification, chi-square tests for emotion selection and perceived emotional change over time, and paired t-tests for immersion differences.

The physiological component uses a Polar H10 chest strap for heart-rate monitoring. The protocol records HR (bpm) and RR intervals (ms) at 1 Hz. One participant was excluded for declining the sensor, incomplete recordings were removed, and implausible values were excluded using the thresholds HR: 46–200 bpm and RR: 300–1300 ms. Missing values were interpolated using Piecewise Cubic Hermite Interpolating Polynomial (PCHIP), and windows with more than 30% interpolated samples were excluded. Metrics were then computed in the time domain and frequency domain. The extracted measures were: SDNN, RMSSD, RR Mean, pNN20, pNN50; Mean and Median HR, HR STD, IQR HR, Mean Difference; and VLF, LF, and HF Power, Total Power, and LF/HF.

The behavioural component uses video-based motion tracking with OpenPose via OpenPifPaf from stationary cameras. Only keypoints above a 10% confidence threshold were used. Participants were assigned through manually defined bounding boxes, frames were temporally subsampled to about 1 Hz, low-quality footage was excluded, and total movement was computed from frame-to-frame skeletal keypoint displacements, normalized by bounding-box size and weighted by confidence. The reported derived measures were Mean and SD Movement and Mean and SD Synchrony. Synchrony was defined as pairwise alignment of movement vectors using cosine similarity on a scale from 0 to 1, where 1 indicates strong synchrony and 0 indicates none.

Unlike formal signal-processing studies organized around explicit equations, EMORSION does not provide major LaTeX equations for the audio protocol itself. Its technical specificity instead lies in metric definitions, preprocessing rules, and inferential structure, including Benjamini-Hochberg correction.

5. Reported response patterns

The protocol detected measurable differences across audio conditions, and the reported findings indicate that even subtle changes in audio design can shape emotional perception and immersion (Garcia et al., 29 May 2026).

In the self-report data, the clearest effects appeared in immersion ratings. Most participants reported increased immersion for at least one augmented mix per scene. Directionality had the strongest impact for A Quiet Place (AQP) and Decision to Leave (DTL), whereas frequency altered immersion for Ford vs Ferrari (FVF) and I Saw the TV Glow (ISTVG). Reported statistically significant immersion p-values included FVF: original and frequency, both 0.01; AQP: original and directionality, both 0.002; DTL: original and directionality, both 0.02; and ISTVG: original and dynamics/frequency, with 0.03 and 0.0006. By contrast, intensity-change p-values were not statistically significant.

In the physiological results, frequency-domain HRV measures were mostly non-significant after correction, with one reported exception: a marginal total power effect for the DTL directionality mix at p = 0.035. Time-domain measures were more sensitive, especially for dynamics. SDNN was significant for DTL dynamics (p = 0.014) and AQP dynamics (p < 0.001). HR standard deviation was significant for DTL dynamics (p = 0.035), AQP dynamics (p < 0.001), and DTL frequency (p = 0.035). HR interquartile range was significant for DTL dynamics (p = 0.04), AQP dynamics (p = 0.01), and ISTVG dynamics (p = 0.01). The discussion interprets these results as indicating that dynamics changes seemed most physiologically salient, that AQP showed reduced HRV consistent with sustained arousal, and that DTL and ISTVG showed HR increases, suggesting heightened reactivity.

In the motion-tracking results, overall mean audience movement ranged from 122 to 535, with a global mean = 330. Horror scenesISTVG and AQP—tended to produce lower movement, while FVF and DTL produced higher movement. Directionality mixes increased mean movement relative to original mixes. Synchrony values were generally higher during second viewings, suggesting greater collective alignment over time. FVF showed stable synchrony and low inter-participant variation, while ISTVG showed very low movement but relatively high synchrony, consistent with a more still, collectively absorbed viewing style.

A prominent qualitative conclusion is the contrast between unconventional and conventional immersive mixes. The study reports that unconventional mixes tended to produce greater variability in audience interpretation, whereas conventional immersive mixes were associated with stronger cross-audience agreement. In other words, audio manipulations aligned with expected immersive design elicited more convergent reports of what was heard and felt, while more unusual manipulations produced more dispersed responses.

6. Significance, feasibility, and limitations

EMORSION is significant because it demonstrates that audio parameters can be experimentally manipulated in cinema conditions and that resulting audience responses can be captured through a triangulated multimodal method. The protocol therefore pushes film-sound research beyond music-only studies and beyond highly artificial lab tasks. Its contribution is methodological as much as empirical: it establishes that real-audience measurement across subjective, physiological, and behavioural channels is operationally possible in a functioning cinema environment (Garcia et al., 29 May 2026).

The study presents the protocol as feasible for three reasons. First, the cinema setup worked as an exhibition environment for controlled stimulus presentation. Second, audiences could be measured simultaneously across multiple modalities. Third, the protocol produced interpretable differences across audio conditions, supporting its use as a proof-of-concept for larger investigations.

The limitations are stated explicitly. Movement tracking was the hardest measure to obtain reliably, and one session yielded almost no usable movement data. The 1 fps temporal subsampling limited motion resolution. Limited access to commercial stems constrained mix precision. Sensor availability and camera coverage limited sample size and reduced peripheral tracking accuracy. More generally, the reported effects were complex and scene-dependent rather than uniform across all manipulations.

Accordingly, EMORSION is best understood as a successful proof-of-concept rather than a complete parametric model of film-audio causation. The study argues for larger-scale studies to determine more precisely how specific audio parameters shape narrative engagement, emotional response, and immersion. A plausible implication is that future work could use the protocol to separate scene-specific effects from more general regularities in how dynamics, directionality, and frequency contribute to collective cinematic experience.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to EMORSION Protocol.