First-Person Physical Perspective in Robotics & VR
- First-Person Physical Perspective (FPP) is an embodied, egocentric viewpoint that couples perception and action from a body-centered reference frame.
- FPP is applied in human–robot interaction, immersive VR, and computer vision to enhance self-location, proxemics, and task-specific performance.
- Computational models of FPP integrate multisensory inputs and dynamic reference frames to predict trajectories, social responses, and affordances.
Searching arXiv for papers on first-person perspective, egocentric/embodied viewpoints, and related evaluation frameworks. First-Person Physical Perspective (FPP) denotes an embodied, subjective, egocentric viewpoint in which an agent is physically in a scene and experiences objects, trajectories, and other agents relative to its own body. In human–robot interaction, FPP is the viewpoint of a pedestrian who is physically in the robot’s path rather than an observer looking from above; in immersive virtual reality, it corresponds to first-person perspective (1PP), where the virtual camera is aligned with the user’s own viewpoint and self-location is expected to coincide with the avatar; in egocentric vision, it is the moving camera-centered frame in which future locations, attention, and affordances are predicted (Agrawal et al., 30 Mar 2026, Anjos et al., 2024, Yagi et al., 2017). Across these uses, FPP is not merely a camera placement. It is a body-centered reference frame that couples perception, action, proxemics, and often self-location.
1. Conceptual scope and reference frames
A recurring contrast in the literature is between allocentric and egocentric description. The allocentric view is a bird’s-eye, detached, scene-level representation; the egocentric view is tied to the subject’s bodily location. In the robot-sociability study, three perspectives are explicitly defined: allocentric (2D overhead), egocentric-proximal (first-person view from the pedestrian closest to the robot’s path), and egocentric-distal (first-person view from the farthest pedestrian in the group). The paper identifies the two egocentric conditions as FPP and states that allocentric perspective is not FPP because it is abstract, detached, and emphasizes global motion patterns rather than local bodily impact (Agrawal et al., 30 Mar 2026).
In immersive VR, the same distinction appears as first-person versus third-person embodiment. First-Person Perspective (1PP) is the egocentric viewpoint with the virtual camera placed at the user’s own viewpoint and aligned with the head-mounted display; Third-Person Perspective (3PP) uses a displaced camera behind and above the user, producing an “out-of-body experience” in which body ownership may persist but self-location becomes displaced or ambiguous (Anjos et al., 2024). In teleoperation with supernumerary limbs, the closest analogue to FPP is the Shared Embodied View, where the guest sees from the host’s head position and shares the host’s bodily frame while controlling extra arms (Zhou et al., 31 Jan 2026).
| Perspective form | Description in the literature | Relation to FPP |
|---|---|---|
| Allocentric | 2D overhead, bird’s-eye, detached scene view | Not FPP |
| Egocentric-proximal | First-person view from the pedestrian closest to the robot’s path | Strongest FPP |
| Egocentric-distal | First-person view from the farthest pedestrian in the group | FPP with less physical intrusion |
| 1PP | Virtual camera aligned with the HMD and user viewpoint | FPP in immersive VR |
| 3PP | Camera displaced behind and above the user | Displaced from FPP |
This framing matters because many results in the literature arise not from different physical trajectories but from different reference frames imposed on the same underlying scene. A central empirical claim is therefore that perspective is itself a causal variable in evaluation, not merely a visualization choice (Agrawal et al., 30 Mar 2026).
2. Embodiment, self-location, and viewpoint control
The immersive-VR literature treats FPP as a structured embodiment problem. The “sense of embodiment,” following Kilteni et al. as summarized in the avatar study, comprises agency, body ownership, and self-location. In 1PP, self-location is expected to coincide with the avatar; in 3PP, self-location is visually displaced, and the virtual head is not exactly where the user’s real head is (Anjos et al., 2024). Empirically, self-location scores were consistently higher in 1PP than in 3PP: for Q3, the medians were $5(2)$, $5(3)$, and $4(2)$ in 1PP for abstract, mesh, and point-cloud avatars, versus $3(3)$, $2.5(2)$, and $3(3)$ in 3PP (Anjos et al., 2024).
These embodiment differences produce task-specific trade-offs rather than a universal first-person advantage. In the same study, 1PP was faster than 3PP for all avatars on Task 1 (horizontal navigation) and Task 3 (ducking under a bar), and reflex-based tasks had better performance in 1PP. Yet 3PP was recommended when spatial awareness on the horizontal navigation plane was required, with a slight advantage when using a realistic representation (Anjos et al., 2024). The paper further concluded that the uncanny valley effect is more prevalent in the third person perspective, whereas in 1PP it is mainly noticed in tasks where a higher sense of embodiment is required.
Teleoperation studies reach a similar conclusion from a different direction. In “One Body, Two Minds,” Shared Embodied View is the FPP-like baseline, but Embedded Anchored View supports embodiment while Out-of-body View improves navigation efficiency and reduces errors. The authors therefore recommend Embedded Anchored View for hand-centric adjustments and Out-of-body View for navigation and object placement, with smooth transitions between them (Zhou et al., 31 Jan 2026). In UAV control through a digital twin, First-Person View was implemented by co-locating the VR camera with the physical drone. Quantitative analysis showed that First-Person View was associated with significantly higher mental demand and effort, greater trajectory deviation, but smoother control inputs than Third-Person and Chase perspectives, and preference data indicated that Third-Person View was most consistently favored whereas First-Person View elicited polarized reactions (Vona et al., 22 Oct 2025). This suggests that FPP strengthens self-location and local control continuity, but can simultaneously raise workload when global spatial updating is difficult.
3. Proxemics, sociability, comfort, and the perceptual gap
The most explicit FPP result in human–robot interaction is the claim that allocentric validation systematically underestimates discomfort and overestimates sociability compared to what people feel from FPP, especially in the egocentric-proximal case (Agrawal et al., 30 Mar 2026). In the VR-based user study, the physical scene and robot trajectories were identical across perspectives; only the camera viewpoint changed. The proximal first-person condition placed the robot at a minimum distance of about , whereas the distal first-person condition used approximately . This design isolated perspective as the causal factor in changes to perceived sociability and disturbance (Agrawal et al., 30 Mar 2026).
The quantitative results establish a proxemic asymmetry. Disturbance showed a significant main effect of condition, , and proximal FPP was significantly more disturbing than distal FPP and more disturbing than allocentric perspective $5(3)$0 (Agrawal et al., 30 Mar 2026). Sociability also varied significantly across conditions, $5(3)$1, with distal FPP significantly more sociable than proximal FPP $5(3)$2, even though allocentric ratings were not significantly different from either egocentric condition on that subscale (Agrawal et al., 30 Mar 2026). The paper’s central inference is that allocentric “OK” trajectories can hide a first-person discomfort that appears only when a person actually stands in the robot’s path.
The same study also separates physical intrusion from social signaling. Augmenting the RL-based trajectory with a head nod increased sociability in egocentric-proximal FPP $5(3)$3 and egocentric-distal FPP $5(3)$4, but had no significant effect on disturbance in either condition (Agrawal et al., 30 Mar 2026). The implication made explicit in the paper is that spacing is primary for disturbance, whereas signaling modulates positive appraisal.
A related, but distinct, first-person result appears in immersive telepresence. In “Defining Preferred and Natural Robot Motions in Immersive Telepresence from a First-Person Perspective,” comfort had a strong effect on path preference, and the subjective feeling of naturalness also had a strong effect on path preference, even though people considered different things natural (Mimnaugh et al., 2021). Participants preferred paths with forward speed of $5(3)$5, focused strongly on turns and on distance to walls and salient objects, and more often preferred piecewise linear paths with in-place rotations over the smooth RRT-based path (Mimnaugh et al., 2021). This broadens the FPP problem from social proxemics to trajectory phenomenology: first-person experience is shaped not only by clearance and collision risk, but by how turning, speed, and route structure feel when the robot’s body is perceptually one’s own.
4. Computational formalizations of FPP
In computer vision, FPP is operationalized as a moving egocentric frame whose geometry must be modeled explicitly. “Future Person Localization in First-Person Videos” formalizes person prediction in this frame by representing a target person’s image-plane location as $5(3)$6, scale as $5(3)$7, and the joint location-scale vector as $5(3)$8 (Yagi et al., 2017). Ego-motion is encoded either as cumulative 3D rotation–translation features $5(3)$9 or, in the Social Interaction Dataset, as a 24D optical-flow descriptor. These streams, together with a pose stream $4(2)$0, are fused in a multi-stream convolution–deconvolution network trained with mean squared error on future relative location-scale trajectories. The full model achieved average FDE $4(2)$1 pixels on the First-Person Locomotion dataset, compared with $4(2)$2 for ConstVel, $4(2)$3 for NNeighbor, and $4(2)$4 for Social LSTM, showing that ego-motion, scale, and pose are not auxiliary cues but constitutive variables of egocentric prediction (Yagi et al., 2017).
Egocentric field-of-view localization uses a complementary formulation. “Egocentric Field-of-View Localization Using First-Person Point-of-View Devices” defines the focus of attention in the POV image as $4(2)$5, transfers it to a reference image through an affine transformation $4(2)$6, and fuses the visually estimated point $4(2)$7 with a sensor-based head-orientation estimate $4(2)$8 through $4(2)$9 (Bettadapura et al., 2015). On outdoor data, accuracy improved from $3(3)$0 without sensor data to $3(3)$1 with sensor fusion; on the indoor presentation task, FOV localization and camera selection were correct in $3(3)$2 cases, or $3(3)$3; on museum images, localization within the correct painting frame reached $3(3)$4 (Bettadapura et al., 2015). Here FPP is the jointly visual and inertial estimate of where a person is and what that person is attending to in a shared environment.
The same agent-centered logic appears in person re-identification. EgoReID contributes a dataset captured using three mobile phones with non-overlapping FOV, containing 900 IDs, around 10,200 tracks, about 176,000 detections, and 12-sensor metadata including camera orientation pitch and rotation. Its method combines semantic body parsing with 3D temporal convolutions, then uses GPS, heading, and speed metadata to predict the target’s next camera and estimated time of arrival, reducing the ReID search space (Basaran et al., 2018). On EgoReID, the visual model with full, upper, lower, and global features reached Rank-1 $3(3)$5 and mAP $3(3)$6; adding metadata increased these to Rank-1 $3(3)$7 and mAP $3(3)$8 (Basaran et al., 2018). This is a canonical FPP effect: once the observer is mobile, the camera’s physical state becomes part of the identity-matching problem.
A more speculative computational use of FPP appears in inverse visual path planning. “Customizing First Person Image Through Desired Actions” states the conjecture that the spatial arrangement of a first person visual scene is deployed to afford an action, and that the action can be inversely used to synthesize a new scene such that the action is feasible. Its key construct is ActionTunnel, a 3D virtual tunnel along the future trajectory encoding what the wearer will visually experience as moving into the scene (Su et al., 2017). Even from the abstract alone, this places FPP within an affordance-based loop in which future bodily motion constrains what a first-person scene should look like.
5. Perspective transformation, shared cognition, and first-person AI
A second line of work treats FPP not only as a viewpoint but as a target representation that can be generated, shared, or benchmarked. In “Diffusing in Someone Else’s Shoes,” a conditional diffusion model learns to generate a robot’s first-person perspective from third-person demonstrations, either as first-person RGB images or as joint vectors corresponding to that egocentric view (Spisak et al., 2024). The model directly predicts the denoised target $3(3)$9 rather than noise, using a U-Net-style architecture with a separate condition path for the third-person image. Quantitatively, the image-to-image perspective transform outperformed CycleGAN and Pix2pix with MSE $2.5(2)$0, L1 $2.5(2)$1, and SSIM $2.5(2)$2, versus CycleGAN’s $2.5(2)$3, $2.5(2)$4, $2.5(2)$5 and Pix2pix’s $2.5(2)$6, $2.5(2)$7, $2.5(2)$8 (Spisak et al., 2024). FPP is therefore modeled as an embodied data space that can be inferred from exocentric demonstrations.
In human–AI collaboration, FPP becomes a shared perception channel. Eye2Eye uses the first-person perspective captured by Apple Vision Pro to coordinate joint attention, revisable memory, and reflective feedback, with object-card retrieval defined by cosine similarity and cumulative alignment tracked by a cumulative error rate over turns (Teng et al., 13 Mar 2026). In the user study, Eye2Eye reduced error rate from EMM $2.5(2)$9 to $3(3)$0, reduced clarification cost from $3(3)$1 turns to $3(3)$2, and improved fluency, copresence, shared awareness, and performance trust (Teng et al., 13 Mar 2026). This result is important because it moves FPP beyond viewpoint realism: the shared egocentric frame becomes the substrate for common ground.
Benchmark work extends the same idea into model evaluation. EgoSocialArena argues that first-person, ego-centric evaluation aligns better with actual LLM-based agent use scenarios than third-person social reasoning, and reports that even the most advanced LLMs lag behind human performance in several first-person social intelligence tasks (Hou et al., 2024). EgoThink evaluates vision–LLMs on twelve first-person dimensions built from egocentric images, including existence, affordance, spatial relationship, situated reasoning, forecasting, navigation, and assistance (Cheng et al., 2023). GPT-4V led overall with average score $3(3)$3, but performed only $3(3)$4 on counting and $3(3)$5 on forecasting, while many open models remained well below that level (Cheng et al., 2023). These benchmarks make explicit that FPP is not exhausted by geometric alignment; it also requires self-centered inference about action, affordance, and future events.
6. Philosophical extensions, misconceptions, and open problems
The widest extension of FPP appears in philosophy of physics. “On Participatory Realism” argues that quantum theory exerts a “nagging pressure to insert a first-person perspective into the heart of physics,” and that reality is more than any third-person perspective can capture (Fuchs, 2016). In QBism, the Born rule is written as a normative relation among an agent’s personal probabilities,
$3(3)$6
so that the “I” is not eliminated from physics but made explicit as an acting, experiencing agent (Fuchs, 2016). This use is not empirical FPP in the HRI or VR sense, but it highlights a common conceptual thread: first-person perspective is treated as constitutive rather than incidental.
Several misconceptions recur across the empirical literature. One is that an egocentric camera alone suffices for FPP. The studies do not support that simplification. Disturbance in robot navigation depends on bodily proximity and peripersonal intrusion, not merely viewpoint; a direct HMD feed can provide first-person visual perspective while still hiding bodily interaction; and 1PP improves self-location without always improving global navigation or obstacle avoidance (Agrawal et al., 30 Mar 2026, Emmerich et al., 2021, Anjos et al., 2024). A second misconception is that smoother or more “human-like” trajectories are automatically preferable from within the moving body. In telepresence, piecewise linear paths with in-place rotations were often preferred over a smooth RRT path, and in UAV control the first-person drone view produced smoother control inputs but worse trajectory deviation and higher workload than external views (Mimnaugh et al., 2021, Vona et al., 22 Oct 2025). A third misconception is that social signaling can compensate for physical intrusion. The head-nod study shows the opposite pattern: signaling raised sociability but did not significantly reduce disturbance when the robot passed too close (Agrawal et al., 30 Mar 2026).
Open problems follow directly from these tensions. The robot-sociability study points to systematic variation of passing distance and approach angle to derive quantitative FPP comfort curves, and to incorporating FPP-based feedback into RL reward functions (Agrawal et al., 30 Mar 2026). The VR embodiment study identifies unresolved questions about vertical-plane obstacle avoidance, poor estimation of distance, and the effect of body representation in social and collaborative tasks (Anjos et al., 2024). Eye2Eye raises issues of latency, thresholded intent triggers, privacy, and the need for in-the-wild studies (Teng et al., 13 Mar 2026). Teleoperation work indicates that perspective switching itself becomes a design variable: FPP-like shared embodiment may be indispensable for near-body manipulation, yet external or stabilized views may be better for navigation and object placement (Zhou et al., 31 Jan 2026). This suggests that future FPP research will likely treat first-person perspective not as a single mode to be maximized, but as one component in adaptive, task-dependent systems that negotiate among embodiment, situational awareness, comfort, and interpretability.