---
title: 'Eyes-on-Me: Gaze & Attention in Interaction'
url: https://www.emergentmind.com/topics/eyes-on-me
type: topic
---

# Eyes-on-Me: Gaze & Attention in Interaction

“Eyes-on-Me” denotes a family of research problems and interface designs centered on directed attention: whether a human, avatar, or robot can establish, detect, display, or exploit gaze in a way that makes attention legible to others. In recent work, the phrase is most closely associated with mutual gaze awareness in telepresence and human–AI communication, but it also appears in gaze-triggered activation, eye-contact coaching, robot explainability, video-based inference of interpersonal gaze, and, in a separate metaphorical sense, attention steering in retrieval-augmented generation (RAG) systems [2407.05833][2510.00586].

## 1. Mutual gaze as the core meaning

Within mediated communication, the central meaning of Eyes-on-Me is full mutual gaze rather than unidirectional gaze awareness. “See-Through Face Display” presents an eye-contact display system designed to enhance gaze awareness in both human-to-human and human-to-avatar communication, and its governing principle is “what you see is what you show”: if the user looks at the interlocutor’s eyes on the display, their own camera view shows mutual eye contact. Earlier “face display” technology allowed only unidirectional gaze awareness—the remote user cannot know if they are being looked at—whereas See-Through Face Display achieves full mutual gaze; both parties know they are being looked at [2407.05833].

The same work ties mutual gaze to social presence rather than to eye geometry alone. In a 3-person teleconference, when two people look at each other, they perceive genuine eye contact, closely resembling real face-to-face interaction, while the third person observes the dyad making eye contact and feels excluded. The reported qualitative finding is explicit: “Users unable to make eye contact reported a sense of exclusion similar to being left out in a face-to-face conversation.” Users also felt as if remote partners were “physically present.” Captured video and audio are sent via WebRTC, and multiple displays can be arranged for multi-party calls, each with its own face video feed, supporting parallel mutual and triadic gaze [2407.05833].

This usage makes Eyes-on-Me a property of an interaction loop. The relevant question is not simply whether gaze is estimated, but whether all participants can perceive who is looking at whom, and whether that perception reproduces familiar conversational effects such as engagement, turn-taking, and exclusion.

## 2. Optical alignment and behind-display architectures

The most direct hardware solution to Eyes-on-Me is physical alignment between display, camera, and perceived eye position. In See-Through Face Display, the hardware is a 4-inch, full-color transparent LCD (Japan Display Inc.), 320×360 pixels, with Toshiba Teli’s BU160MCF camera placed directly behind the display. At 60 FPS, i.e. every 16.7 ms, the display is transparent for 6 ms, and the camera captures precisely during those intervals. This allows the display to show the remote user’s face while the camera directly captures the user’s actual gaze from the correct line of sight, eliminating parallax errors common to traditional videoconferencing setups [2407.05833].

“AnimeGaze” extends the same camera-behind-the-display principle from human face video to rendered avatars in physical environments. Its display alternates at 50 Hz refresh and 180 Hz field sequential, using a 4" diagonal, 320×360 display paired with a 1440×1080 camera. The key geometric statement is that the avatar’s virtual eyes are co-located with the camera center, $\mathbf{A}=\mathbf{C}$, so the camera ray to a real-world target and the avatar’s gaze ray are identical. This enables two-way gaze communication with people and objects in the physical environment, supports avatars with any number of eyes, and includes calibration to reduce the Mona Lisa effect through a mapping between actual and perceived gaze directions [2503.06324].

A related alignment problem appears in optical see-through AR, where the world-facing camera and the user’s eye perspective are inherently misregistered. “See What I Mean?” evaluates Plane-Proxy EPR, Mesh-Proxy EPR, and Gaze-Proxy EPR on Microsoft HoloLens 2. In the reported user study, MESH mean error is 1.1mm and GA mean error is 1.3mm at 75cm viewing distance, while participant rankings favor MESH and GA over PLA. The result is a useful corrective to a common misconception: accurate eyes-on-target perception in AR is not guaranteed by merely projecting camera output onto a fixed plane; the projection must be rendered in the user’s effective eye perspective [2509.11653].

Across these systems, the technical objective is the same: reduce the discrepancy between where attention is displayed and where attention is perceived.

## 3. Gaze as an input channel and a trainable skill

A second line of work operationalizes Eyes-on-Me as an interaction trigger. “Look and Talk” uses Microsoft HoloLens 2 with integrated real-time eye-tracking to activate an AI assistant when a user fixates on a virtual avatar for a threshold period. The avatar is rendered slightly above the user’s direct line of sight, and a dwell time threshold $T_{\text{thresh}}$ of 2 seconds is empirically chosen through informal testing to filter out accidental or brief glances. In the reported Wizard of Oz study, fixation-based activation was described as “intuitive and natural,” participants appreciated “no need for wake words,” and they requested stronger activation feedback while noting that 2 seconds felt slightly long [2504.09296].

Here, Eyes-on-Me is not mutual gaze in the telepresence sense; it is intention recognition. Sustained fixation is treated as a socially legible, silent, and hands-free activation signal. That framing is consistent with the paper’s claim that the method is inspired by human conversational cues and fits seamlessly into natural human behavior [2504.09296].

Public speaking assistance treats the same phenomenon as a skill that can be monitored and corrected in real time. “Talk to Me, Not the Slides” presents SpeakAssis, a wearable system based on the Pupil Core eye tracker, with a front-facing scene camera at 1280×720 and two near-eye cameras at 120 Hz. Its anchor face tracking design means that only 6% of frames require heavy face identification; the remainder use lightweight tracking. The reported performance is 93.3% face identification accuracy with 28 ms average latency per frame, compared to 30.8% accuracy and 57 ms latency for the baseline. In the user study, SpeakAssis increases speakers’ eye-contact duration by 62.5% on average and yields an average 17.4% increase in Gaze Distribution Entropy. Audience surveys further report that improvements in speaker’s eye-contact behavior significantly enhance perceived engagement and interactivity, although audio prompts can be intrusive [2602.01201].

The technical significance of these systems is that gaze becomes both an implicit command channel and a target of feedback control. Eyes-on-Me is therefore not only a state to be detected; it is a behavior to be elicited, tuned, and redistributed.

## 4. Explainability, joint attention, and synthetic gaze

In human–robot interaction, Eyes-on-Me often denotes explainability through attention display. “Mirror Eyes” equips a robot head with screen-based eyes that can direct gaze to points in physical space and overlay a live, horizontally flipped image segment from the attended region onto each pupil. In a user study with 33 participants performing pick-and-place supervision, Mirror Eyes improves Subjective Information Processing Awareness from 3.42 (SD = 1.16) to 4.9 (SD = 1.01), with \(t(32)=8.17, p<.001\). Participants also stop erroneous actions faster: early-error interruption time decreases from 5.52s (SD = 2.77) to 4.66s (SD = 1.85), and late-error interruption time from 15.76s (SD = 1.31) to 14.58s (SD = 2.9). Pragmatic, hedonic, and overall UEQ-S scores are all higher with \(p<.001\). The benefits appear even without instruction regarding the eyes, and the method is explicitly described as supporting the ‘Eyes-on-Me’ phenomenon by drawing and signaling attention, reference, and communicative intention through gaze and mirrored attended regions [2506.18466].

This makes a useful distinction between gaze as social contact and gaze as explanatory interface. The robot is not merely “looking at” a target; it is revealing the target of its information processing. A plausible implication is that Eyes-on-Me can function as a transparency mechanism in cooperative systems, especially where rapid human intervention is needed.

Synthetic avatars introduce the complementary problem of generating believable gaze when direct sensing is unavailable or undesirable. “TalkingEyes” constructs the 3D Talking Eyes Dataset with 5,982 videos, 669 subjects, and about 14 hours of footage, and models head motion and eye gaze motion from speech in separate latent spaces. On its quantitative evaluation, TalkingEyes reports Diversity \(0.264\), Corr. \(0.604\), and L2 \(0.266\), while user preference rates are 83% or above [2501.09921].

The Eyes-on-Me problem in this setting is not optical alignment but behavioral plausibility. Speech-driven gaze animation addresses whether an avatar can appear attentive and socially responsive even when gaze is synthesized rather than captured.

## 5. Measurement and inference: from onfocus to interpersonal gaze

A persistent source of confusion in this area is the difference between individual-camera eye contact and interpersonal gaze. “Onfocus Detection” defines onfocus detection as determining whether the focus or gaze of an individual in an unconstrained image is directed at the capturing camera, i.e., whether a person or animal is making eye contact with the camera. Its OFDIW dataset contains 20,623 images, and the Eye-Context Interaction Inferring Network (ECIIN) obtains Accuracy 0.8471 and F-measure 0.9007 on OFDIW-HF; on OFDIW-PF, ECIIN obtains 0.7744 and 0.7747 [2103.15307].

By contrast, “Are you really looking at me?” targets interpersonal eye gaze relative to a conversation partner rather than to the camera. The Interpersonal-Calibrating Eye-gaze Encoder (ICE) automatically extracts interpersonal gaze from video recordings without specialized hardware and without prior knowledge of participant locations. It is validated in video chat with an infrared gaze tracker at F1 = 0.846, \(N=8\), and in face-to-face communication with expert-rated evaluations of eye contact at \(r=0.37\), \(N=170\). The behavioral findings are consequential: honest witnesses break interpersonal gaze contact and look down more often than deceptive witnesses when answering questions (\(p=0.004, d=0.79\)), and interpersonal gaze alone has more predictive power than facial expressions in predicting expert communication skill ratings in speed dating videos [1906.12175].

Lower-level gaze estimation methods supply the estimation substrate for both onfocus and interpersonal-gaze systems. “End-to-end Video Gaze Estimation via Capturing Head-face-eye Spatial-temporal Interaction Context” reports a mean angular error of 10.02° on the Gaze360 detectable face subset, 12.96° on the entire Gaze360, and 70 FPS for clip length 7 on a single RTX 3090, while requiring no external detections or pre-processing [2310.18131].

Taken together, these works show that Eyes-on-Me can be formalized at several observational levels: direct camera eye contact, partner-relative gaze, and continuous 3D gaze regression. The technical boundary between them is important, because each task encodes a different social referent.

## 6. Metaphorical extension: attention steering in machine learning systems

Outside literal eye contact, the phrase has also been adopted to describe attention steering inside machine learning systems. “Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors” decomposes an adversarial document into reusable Attention Attractors and Focus Regions, optimizes a small subset of attention heads that are strongly correlated with attack success, and uses the resulting attractors to steer retrieval and generation toward malicious payloads. Across 18 end-to-end RAG settings, the method raises average attack success rates from 21.9 to 57.8 (+35.9 points, 2.6× over prior work). The threat model uses one poisoned document per 1,000-document corpus (<0.1%), and a single optimized attractor transfers to unseen black box retrievers and generators without retraining [2510.00586].

This is a metaphorical, not social, use of Eyes-on-Me. The “eyes” are attention heads rather than human observers, and the objective is not mutual gaze awareness but reliable concentration of model attention on a focus region. Even so, the terminological continuity is revealing. This suggests that the phrase now functions as a compact label for systems in which an internal locus of attention is deliberately aligned with an externally consequential target, whether that target is a remote interlocutor, a robot’s attended object, an audience segment, or a poisoned document region.

Source: https://www.emergentmind.com/topics/eyes-on-me