Papers
Topics
Authors
Recent
Search
2000 character limit reached

VRSL:Exploring the Comprehensibility of 360-Degree Camera Feeds for Sign Language Communication in Virtual Reality

Published 26 Feb 2026 in cs.HC | (2602.23265v1)

Abstract: This study explores integrating sign language into virtual reality (VR) by examining the comprehensibility and user experience of viewing American Sign Language (ASL) videos captured with body-mounted 360-degree cameras. Ten participants identified ASL signs from videos recorded at three body-mounted positions: head, shoulder, and chest. Results showed the shoulder-mounted camera achieved the highest accuracy (85%), though differences between positions were not statistically significant. Participants noted that peripheral distortion in 360-degree videos impacted clarity, highlighting areas for improvement. Despite challenges, the overall comprehension success rate of 83.3% demonstrates the potential of video-based ASL communication in VR. Feedback emphasized the need to refine camera angles, reduce distortion, and explore alternative mounting positions. Participants expressed a preference for signing over text-based communication in VR, highlighting the importance of developing this approach to enhance accessibility and collaboration for Deaf and Hard of Hearing (DHH) users in virtual environments.

Summary

  • The paper evaluates body-mounted 360-degree cameras for ASL communication in VR through a within-subjects study of 10 participants across head, shoulder, and chest mounts, achieving 83.3% overall comprehension.
  • The shoulder mount produced the highest accuracy at 85%, but differences between camera positions were not statistically significant, indicating that all three configurations may support understandable signing.
  • The results show that location and movement signs were easiest to recognize, while palm orientation and sentence-level interpretation were harder, highlighting the need to reduce optical distortion and test real-time communication.

Motivation and research gap

VR communication relies heavily on auditory channels, which disadvantages Deaf and Hard of Hearing (DHH) users. Existing accessibility mechanisms each fall short for DHH-to-DHH interaction: closed captioning primarily bridges hearing–DHH communication; hand-tracking and avatar-based signing suffer from headset blind spots, imprecise capture of signs involving the face or body, and resource-heavy real-time avatar rendering — a problem visible in VRChat, where users invented ad hoc signs due to poor rendering fidelity; and text chat strips away the non-verbal channel that carries much of sign language's expressive power. A 2022 systematic review of immersive social VR communication identified only one study addressing DHH users among 32 analyzed papers, underscoring how little is known about DHH-to-DHH communication in VR.

The paper proposes an alternative: capturing American Sign Language (ASL) with body-mounted 360-degree cameras and presenting the video feed in VR. The approach targets the five linguistic parameters of ASL — hand shape, palm orientation, movement, location, and non-manual signals — which prior linguistics work identifies as essential to sign comprehension. It builds on "Chat in the Hat," which demonstrated feasibility of body-mounted cameras for interpreted ASL communication between a DHH user and hearing interlocutors.

Study design

The authors conducted a within-subjects study with one independent variable: camera mounting position on the signer (head, shoulder, chest). Ten participants (9 DHH, 1 hearing but ASL-fluent; ages 18–55) viewed pre-recorded 360-degree videos — captured with a Ricoh Theta V, converted to equirectangular format, and rendered in a Unity application on a Meta Quest — and selected the correct interpretation from multiple-choice options, with a skip option to discourage guessing. The signer wore a VR headset during recording to simulate a VR player signing.

Tasks were organized into three categories grounded in ASL linguistics: minimal pairs (signs differing in exactly one parameter), non-minimal pairs (isolated words approximating everyday use), and sentences (testing contextual interpretation). Each condition contained 10 tasks balanced across categories (5 minimal-pair, 2 single-word, 3 sentence tasks), randomized for order. Perceived workload was measured with NASA TLX, and qualitative feedback was collected via per-condition surveys and an exit interview.

Results

Overall comprehension was high: participants correctly identified an average of 25 of 30 tasks, an 83.3% success rate across all conditions. The shoulder-mounted camera achieved the highest accuracy at 85%, with head at 84% and chest lower; however, a repeated-measures ANOVA found no statistically significant differences between positions (P>0.05P>0.05), implying all three mounts are potentially viable. TLX scores marginally favored the shoulder mount, and stated preferences split evenly (4 chest, 4 shoulder, 2 no preference).

Parameter-level analysis of minimal pairs revealed a clear gradient:

Minimal pair type Accuracy
Location 100%
Movement 96.67%
Handshape 90%
Expression 90%
Palm/orientation 80%

Single-word identification reached 96.67%, while sentence interpretation dropped to 70%, indicating that sequential recognition imposes substantially higher cognitive demand than isolated-sign recognition. Notably, six of ten participants preferred signing via this method over texting or chatting in VR — a direct endorsement of video-based communication over the prevailing text alternative.

Discussion

The results address the paper's three research questions as follows. For RQ1, the 83.3% overall success rate demonstrates that body-mounted 360-degree video can convey ASL legibly in VR, though not yet at the reliability needed for fluent conversation. For RQ2, the shoulder mount performed best numerically, plausibly because it approximates a natural third-person viewing angle and minimizes fisheye distortion relative to the chest position; the lack of statistical significance, however, means this advantage should be treated as suggestive rather than established. For RQ3, participant feedback identified three principal limitations of the method relative to existing tools: peripheral distortion from the dual-fisheye capture, which degraded fine finger articulation critical for minimal pairs; occlusion of signs performed on the side opposite the camera in the shoulder condition ("I only caught half of it"); and unnatural viewing angles, particularly the head-mounted bird's-eye perspective, which diverges from the face-to-face geometry of natural signing.

The parameter-level findings carry design implications: large-scale spatial movements (location, movement) survive 360-degree compression well, whereas fine-grained hand rotations (palm orientation) are most vulnerable to distortion — suggesting that any refinement of this pipeline should prioritize resolution and optics near the hands rather than uniform image quality.

Limitations and open questions

The authors are explicit about several constraints. The study used pre-recorded videos rather than live streams, so real-time latency, bandwidth, and collaboration dynamics remain untested; the reported accuracies therefore bound what a live system could achieve rather than demonstrate one. The sample of ten participants from a single institution is small, and the absence of statistically significant differences between conditions may partly reflect low statistical power. Distortion inherent to fisheye 360-degree capture degraded peripheral sign clarity, motivating exploration of non-fisheye or multi-camera rigs and possibly non-body-mounted cameras offering a more natural third-person view. The shoulder mount's handedness bias — favoring the dominant-hand side where the camera sits — remains unresolved. Open questions include whether alternative mounting positions or external cameras can eliminate the occlusion and angle problems, and how multimodal captions might complement video feeds in mixed-proficiency groups.

Conclusion

This study establishes a preliminary but quantified case for body-mounted 360-degree video as a medium for ASL communication in VR: an 83.3% overall comprehension rate, near-ceiling performance on location and movement parameters, and a majority preference for signing over text chat among DHH participants. The absence of significant differences across head, shoulder, and chest mounts suggests flexibility in hardware placement, while the 70% sentence-level accuracy and palm-orientation errors delineate precisely where current capture quality falls short. Progress toward practical deployment hinges on reducing optical distortion, resolving unilateral occlusion, and validating the approach under real-time, multi-party conditions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.