- The paper evaluates body-mounted 360-degree cameras for ASL communication in VR through a within-subjects study of 10 participants across head, shoulder, and chest mounts, achieving 83.3% overall comprehension.
- The shoulder mount produced the highest accuracy at 85%, but differences between camera positions were not statistically significant, indicating that all three configurations may support understandable signing.
- The results show that location and movement signs were easiest to recognize, while palm orientation and sentence-level interpretation were harder, highlighting the need to reduce optical distortion and test real-time communication.
Motivation and research gap
VR communication relies heavily on auditory channels, which disadvantages Deaf and Hard of Hearing (DHH) users. Existing accessibility mechanisms each fall short for DHH-to-DHH interaction: closed captioning primarily bridges hearing–DHH communication; hand-tracking and avatar-based signing suffer from headset blind spots, imprecise capture of signs involving the face or body, and resource-heavy real-time avatar rendering — a problem visible in VRChat, where users invented ad hoc signs due to poor rendering fidelity; and text chat strips away the non-verbal channel that carries much of sign language's expressive power. A 2022 systematic review of immersive social VR communication identified only one study addressing DHH users among 32 analyzed papers, underscoring how little is known about DHH-to-DHH communication in VR.
The paper proposes an alternative: capturing American Sign Language (ASL) with body-mounted 360-degree cameras and presenting the video feed in VR. The approach targets the five linguistic parameters of ASL — hand shape, palm orientation, movement, location, and non-manual signals — which prior linguistics work identifies as essential to sign comprehension. It builds on "Chat in the Hat," which demonstrated feasibility of body-mounted cameras for interpreted ASL communication between a DHH user and hearing interlocutors.
Study design
The authors conducted a within-subjects study with one independent variable: camera mounting position on the signer (head, shoulder, chest). Ten participants (9 DHH, 1 hearing but ASL-fluent; ages 18–55) viewed pre-recorded 360-degree videos — captured with a Ricoh Theta V, converted to equirectangular format, and rendered in a Unity application on a Meta Quest — and selected the correct interpretation from multiple-choice options, with a skip option to discourage guessing. The signer wore a VR headset during recording to simulate a VR player signing.
Tasks were organized into three categories grounded in ASL linguistics: minimal pairs (signs differing in exactly one parameter), non-minimal pairs (isolated words approximating everyday use), and sentences (testing contextual interpretation). Each condition contained 10 tasks balanced across categories (5 minimal-pair, 2 single-word, 3 sentence tasks), randomized for order. Perceived workload was measured with NASA TLX, and qualitative feedback was collected via per-condition surveys and an exit interview.
Results
Overall comprehension was high: participants correctly identified an average of 25 of 30 tasks, an 83.3% success rate across all conditions. The shoulder-mounted camera achieved the highest accuracy at 85%, with head at 84% and chest lower; however, a repeated-measures ANOVA found no statistically significant differences between positions (P>0.05), implying all three mounts are potentially viable. TLX scores marginally favored the shoulder mount, and stated preferences split evenly (4 chest, 4 shoulder, 2 no preference).
Parameter-level analysis of minimal pairs revealed a clear gradient:
| Minimal pair type |
Accuracy |
| Location |
100% |
| Movement |
96.67% |
| Handshape |
90% |
| Expression |
90% |
| Palm/orientation |
80% |
Single-word identification reached 96.67%, while sentence interpretation dropped to 70%, indicating that sequential recognition imposes substantially higher cognitive demand than isolated-sign recognition. Notably, six of ten participants preferred signing via this method over texting or chatting in VR — a direct endorsement of video-based communication over the prevailing text alternative.
Discussion
The results address the paper's three research questions as follows. For RQ1, the 83.3% overall success rate demonstrates that body-mounted 360-degree video can convey ASL legibly in VR, though not yet at the reliability needed for fluent conversation. For RQ2, the shoulder mount performed best numerically, plausibly because it approximates a natural third-person viewing angle and minimizes fisheye distortion relative to the chest position; the lack of statistical significance, however, means this advantage should be treated as suggestive rather than established. For RQ3, participant feedback identified three principal limitations of the method relative to existing tools: peripheral distortion from the dual-fisheye capture, which degraded fine finger articulation critical for minimal pairs; occlusion of signs performed on the side opposite the camera in the shoulder condition ("I only caught half of it"); and unnatural viewing angles, particularly the head-mounted bird's-eye perspective, which diverges from the face-to-face geometry of natural signing.
The parameter-level findings carry design implications: large-scale spatial movements (location, movement) survive 360-degree compression well, whereas fine-grained hand rotations (palm orientation) are most vulnerable to distortion — suggesting that any refinement of this pipeline should prioritize resolution and optics near the hands rather than uniform image quality.
Limitations and open questions
The authors are explicit about several constraints. The study used pre-recorded videos rather than live streams, so real-time latency, bandwidth, and collaboration dynamics remain untested; the reported accuracies therefore bound what a live system could achieve rather than demonstrate one. The sample of ten participants from a single institution is small, and the absence of statistically significant differences between conditions may partly reflect low statistical power. Distortion inherent to fisheye 360-degree capture degraded peripheral sign clarity, motivating exploration of non-fisheye or multi-camera rigs and possibly non-body-mounted cameras offering a more natural third-person view. The shoulder mount's handedness bias — favoring the dominant-hand side where the camera sits — remains unresolved. Open questions include whether alternative mounting positions or external cameras can eliminate the occlusion and angle problems, and how multimodal captions might complement video feeds in mixed-proficiency groups.
Conclusion
This study establishes a preliminary but quantified case for body-mounted 360-degree video as a medium for ASL communication in VR: an 83.3% overall comprehension rate, near-ceiling performance on location and movement parameters, and a majority preference for signing over text chat among DHH participants. The absence of significant differences across head, shoulder, and chest mounts suggests flexibility in hardware placement, while the 70% sentence-level accuracy and palm-orientation errors delineate precisely where current capture quality falls short. Progress toward practical deployment hinges on reducing optical distortion, resolving unilateral occlusion, and validating the approach under real-time, multi-party conditions.