---
title: Polarization-Enabled Eye Tracking
url: https://www.emergentmind.com/topics/polarization-enabled-eye-tracking-pet
type: topic
---

# Polarization-Enabled Eye Tracking

Searching arXiv for the specified PET papers to ground the article in the cited literature.
Polarization-Enabled Eye Tracking (PET) is a near-infrared eye-tracking paradigm that augments conventional intensity imaging with polarization-resolved measurements of light reflected and scattered from the eye. In the reported formulation, a polarization-filter-array camera paired with a single linearly polarized \(850 \text{ nm}\) illuminator captures four linear-polarization views \((0^\circ,45^\circ,90^\circ,135^\circ)\), and convolutional models use these channels, or derived polarization descriptors, for gaze estimation. The central premise is that polarization introduces an additional optical contrast mechanism: it can reveal dense, trackable scleral texture and repeatable, gaze-informative corneal patterns that are largely absent in intensity-only imagery, thereby broadening the feature set available to end-to-end gaze estimation and reducing brittleness when classic cues such as pupil boundaries and glints degrade [2511.04652].

## 1. Definition and scope

PET denotes an eye-tracking method in which the polarization state of reflected ocular light is measured in addition to its intensity. In the reported systems, the sensing stack comprises a polarization-sensitive camera and linearly polarized NIR illumination; the learned estimator then operates either on the four reconstructed polarization-angle channels or on derived channels such as Intensity, DoLP, and AoLP. This contrasts with standard grayscale eye cameras, which provide one intensity value per pixel, and with RGB imaging, which provides broadband color channels rather than analyzer-angle-resolved measurements [2511.04652].

The motivation for PET arises from known failure modes of appearance-based eye tracking. Standard intensity-only systems rely heavily on pupil boundary localization, corneal glints, and overall limbus or iris appearance. These cues can become unreliable under eyelid or eyelash occlusion, eye-relief changes or slippage, pupil-size variation, and constrained camera placements that provide a poorer view of the eye opening. PET is proposed as a single-camera, single-illuminator alternative to more complex multi-camera architectures by increasing what the source paper calls “input information density” for end-to-end gaze estimation [2511.04652].

A subsequent head-mounted study places PET in a personalization setting. There, the question is not only whether polarization improves generic gaze regression, but whether polarization-sensitive imagery supports low-calibration personalized gaze estimation more effectively than conventional NIR or intensity-only inputs. That study benchmarks PET with a personalized Siamese differential-gaze model and reports that polarization inputs reduce gaze error by up to \(12\%\) compared to intensity-only inputs in the Siamese setting, while 9 calibration anchors achieve performance comparable to linear calibration using about 100 frames [2603.25889].

## 2. Optical basis and polarization representation

In the reported PET systems, the sensor measures reflected light at four linear analyzer orientations:
\[
(0^\circ,\ 45^\circ,\ 90^\circ,\ 135^\circ).
\]
After demosaicking, this yields four per-angle images \(I_{0^\circ}, I_{45^\circ}, I_{90^\circ}, I_{135^\circ}\). From these, the linear Stokes quantities are computed as
\[
S_0 = I_{0^\circ} + I_{45^\circ} + I_{90^\circ} + I_{135^\circ},
\]
\[
S_1 = I_{0^\circ} - I_{90^\circ},
\]
\[
S_2 = I_{45^\circ} - I_{135^\circ}.
\]
Total intensity is then
\[
I = S_0/4,
\]
the degree of linear polarization is
\[
\mathrm{DoLP} = \sqrt{S_1^2 + S_2^2}/(S_0 + \varepsilon),
\]
and the angle of linear polarization is
\[
\mathrm{AoLP} = \tfrac{1}{2}\,\arctan2(S_2, S_1).
\]
Pixels with very low \(S_0\) are masked in DoLP and AoLP maps to suppress artifacts. The reported formulation uses this linear-polarization Stokes representation rather than a full Mueller-matrix model [2511.04652].

The stated optical mechanism is tissue-linked. The sclera and cornea contain birefringent fibrous collagen. In the sclera, anisotropic collagen plus multiple scattering leaves the reflected light with non-zero DoLP and a stable AoLP, producing fine spatial contrast. In the cornea, both interface optics at the air–tear-film–cornea stack and birefringence of corneal stromal lamellae affect the outgoing polarization state. The corneal signal is described as arising from a combination of specular and refracted highlights from layered interfaces and phase retardance or polarization-state rotation introduced by birefringent tissue [2511.04652].

The practical significance of this representation is that regions that are largely featureless in grayscale can become informative in polarization space. The reported qualitative claims are specific: PET reveals dense, fine-grained scleral texture, described as meso-scale collagen “scaffolding,” and repeatable corneal AoLP or DoLP patterns that vary with gaze. This suggests that PET-derived cues are not solely dependent on visible pupil edges or engineered glints, and therefore may remain discriminative when conventional cues are weakened [2511.04652].

## 3. System architectures and processing pipelines

The initial PET demonstration uses a non-form-factor benchtop station with one polarization-filter-array camera per eye and one linearly polarized NIR illuminator per eye. The PET subsystem includes an IDS Imaging UI-3080CP-M-GL Rev.2 camera with a Sony IMX250 PFA sensor, a wire-grid micro-polarizer mosaic, analyzer orientations \((0^\circ,45^\circ,90^\circ,135^\circ)\), and an Osram LZ1-00R402 LED at \(850 \text{ nm}\). A wire-grid linear polarizer film is applied on the illuminator, and the illumination mode is flood illumination rather than structured glint generation. The benchtop system is binocular, and for each eye it evaluates two temporal-side camera positions, designated higher temporal and lower temporal [2511.04652].

Preprocessing in that system is explicit: raw micro-polarizer mosaics are demosaicked into four full-resolution orientation images, Gaussian smoothing with \(\sigma = 1\) is applied, and \(S_0,S_1,S_2\), \(I\), DoLP, and AoLP are computed. For model input, the four polarization channels are used directly and normalized per channel. A lightweight per-user affine correction is learned from a 9-point calibration sequence, consisting of per-eye scale and bias terms on gaze angles; this calibration is then held fixed for robustness tests involving slippage and pupil-size changes [2511.04652].

The end-to-end gaze model in that work is PETNet1. It takes synchronized binocular eye images, processes each eye with a shared-weight 4-stage CNN whose blocks are inverted-residual depthwise-separable \(3 \times 3\) blocks, operates at total stride 16, and outputs a \(25 \times 25\) feature map with 160 channels per eye. Left and right eye features are concatenated to 320 channels for binocular fusion. The backbone has approximately 1.5M parameters; convolutions are bias-free with batch normalization; activations are ReLU; and no cross-view attention or specialized multi-view fusion modules are used. The prediction head maps fused binocular features to per-eye gaze angles. A pseudo-intensity baseline is constructed by averaging the four polarization channels and duplicating that averaged image four times, so that baseline and PET models have matched input dimensionality and matched capacity [2511.04652].

The later head-mounted study adopts a different PET representation and a different learning formulation. It uses a polarization-sensitive camera with \(850 \text{ nm}\) illumination in a binocular head-mounted eye-tracking setup, and its main PET input is a 3-channel tensor comprising Intensity, DoLP, and AoLP. The intensity-only comparison uses three replicated Intensity channels so that both PET and intensity-only inputs are \(3 \times 256 \times 256\). The primary model is a personalized Siamese differential architecture: each branch processes one binocular pair from the same subject, the features from the two branches are concatenated, and a regressor predicts relative gaze displacement \(\Delta g\). Absolute gaze is then reconstructed from anchor images by
\[
\hat{g} = \frac{1}{C}\sum_{c=1}^{C} (\Delta g_c + g_c),
\]
where \(C\) is the number of calibration samples, \(g_c\) is the known gaze target for anchor \(c\), and \(\Delta g_c\) is the predicted displacement between the test input and anchor \(c\) [2603.25889].

## 4. Experimental protocols and quantitative results

The benchtop PET study uses a subject-disjoint split over 346 participants, with 198 training participants and up to 148 validation participants, and trains separate models for each camera placement. The loss is smooth-\(L_1\) (Huber) with outlier rejection, training lasts 400k iterations, and PET and intensity baselines share the same objective and schedule. Data are acquired with a chinrest; eye positions are adjusted to cover a broad distribution of eye relief; and gaze targets are displayed on a monitor at 48 cm distance, with monitor center aligned to \(-9.7^\circ\) tilt relative to nominal \(0^\circ\) gaze and a collection field of view of \(30^\circ \times 20^\circ\). The primary metric is the 95th percentile absolute gaze error per participant, \(E_{95}\), summarized over the population as the median across users, \(U_{50}E_{95}\). Confidence intervals are participant-level nonparametric bootstrapped 90% intervals [2511.04652].

Across all reported conditions and both camera placements, that study reports that PET reduces \(U_{50}E_{95}\) relative to the matched intensity-only baseline by an absolute \(0.12^\circ\) to \(0.23^\circ\) and a relative \(10.3\%\) to \(15.9\%\). The three benchmarked robustness conditions are nominal calibrated operation, eye-relief change without recalibration, and pupil-size change without recalibration. The larger gains in the higher temporal view are explicitly noted as consistent with a more occlusion-prone geometry [2511.04652].

| Condition | Lower temporal | Higher temporal |
|---|---:|---:|
| Nominal calibrated | PET \(1.004^\circ\) vs Intensity \(1.126^\circ\) | PET \(1.185^\circ\) vs Intensity \(1.395^\circ\) |
| Eye-relief change, no recalibration | PET \(1.192^\circ\) vs Intensity \(1.359^\circ\) | PET \(1.551^\circ\) vs Intensity \(1.747^\circ\) |
| Pupil-size change, no recalibration | PET \(1.021^\circ\) vs Intensity \(1.138^\circ\) | PET \(1.199^\circ\) vs Intensity \(1.426^\circ\) |

The same study states that bootstrapped confidence intervals for the PET–intensity median difference exclude zero across broad percentile ranges in all conditions and both camera placements, which it interprets as statistically significant population-level improvements. It also reports a 4-week stability demonstration on one volunteer, imaged on days 1, 5, 7, 15, and 28, where SIFT plus RANSAC on scleral regions yielded 27, 30, 32, and 42 matched inlier keypoints relative to day 1, respectively [2511.04652].

The head-mounted personalization study benchmarks on 338 subjects, with 196 subjects for training and 142 for validation or testing. It reports gaze angular error in degrees at P50, P75, and P95. The main reported comparison is between polarization input and intensity-only input under four settings: Baseline only, Baseline + linear calibration, Siamese only, and Siamese + linear calibration [2603.25889].

| Input and method | P50 | P75 | P95 |
|---|---:|---:|---:|
| Polarization, Siamese + linear calibration | 0.91 | 1.51 | 2.88 |
| Polarization, Siamese only | 1.08 | 1.65 | 2.98 |
| Polarization, Baseline + linear calibration | 1.05 | 1.69 | 3.15 |
| Intensity-only, Siamese only | 1.23 | 1.87 | 3.19 |
| Intensity-only, Baseline + linear calibration | 1.24 | 2.02 | 3.56 |

That study states that 9-anchor Siamese PET achieves \(1.08 / 1.65 / 2.98\), while Baseline + linear calibration with about 100 frames achieves \(1.05 / 1.69 / 3.15\), supporting the claim of comparable performance with about 10-fold fewer calibration samples. It also reports that polarization vs intensity-only in the Siamese setting yields reductions of \(12\%\) at P50, \(12\%\) at P75, and \(6.5\%\) at P95, and that combining Siamese personalization with linear calibration yields further improvements of \(13\%\) at P50, \(11\%\) at P75, and \(8.6\%\) at P95 over a linearly calibrated PET baseline [2603.25889].

## 5. Personalization, calibration, and sample efficiency

Calibration is a central issue in PET because both cited studies frame polarization as a mechanism for improving robustness without eliminating person-specific variation. In the benchtop work, calibration is a lightweight per-user affine correction learned from a 9-point calibration sequence and held fixed during robustness tests involving slippage and pupil-size changes. This choice is important because the reported improvements under non-nominal conditions are measured without recalibration, thereby isolating the contribution of polarization-derived features under degraded operating conditions [2511.04652].

The head-mounted study develops a more explicit personalization framework. Its formal problem defines a binocular input \(I = (I_{left}, I_{right})\) and maps it to
\[
g = [\hat{\theta}_{left}, \hat{\phi}_{left}, \hat{\theta}_{right}, \hat{\phi}_{right}],
\]
with ground truth
\[
g^{gt} = [\theta_{left}, \phi_{left}, \theta_{right}, \phi_{right}].
\]
Instead of directly regressing the final personalized gaze from one sample, the Siamese model learns a differential mapping
\[
f : (I_1, I_2) \to \Delta g
\]
between two binocular eye-image pairs from the same subject. During inference, a small anchor set of calibration images supplies known gaze labels, and final absolute gaze is reconstructed by averaging across predicted displacements relative to those anchors [2603.25889].

The sample-efficiency result is the defining quantitative claim of that study. The number of anchors is varied across 3, 5, 7, and 9. For polarization, the reported Siamese performance improves monotonically with anchor count: \(1.16 / 1.76 / 3.10\) at 3 anchors, \(1.12 / 1.71 / 3.04\) at 5 anchors, \(1.11 / 1.68 / 3.01\) at 7 anchors, and \(1.08 / 1.65 / 2.98\) at 9 anchors. At P95, even 3-anchor Siamese PET is reported to beat the about-100-frame linearly calibrated Baseline PET, \(3.10\) versus \(3.15\). This suggests that polarization cues remain useful even in very low-anchor regimes [2603.25889].

The same study also compares pair-construction strategies for Siamese training. Random same-subject pair sampling achieves \(1.08 / 1.65 / 2.98\) for PET, whereas calibration sampling achieves \(1.34 / 2.05 / 3.56\). The reported interpretation is that random within-subject pairing exposes the model to broader person-specific relative geometry, whereas fixed-anchor pairing encourages memorization. A plausible implication is that PET features are especially compatible with differential learning because the cited scleral and corneal signals are described as both subject-specific and temporally stable [2603.25889].

## 6. Interpretation, limitations, and open directions

Both PET studies interpret the gains as physically grounded rather than merely architectural. The benchtop work argues that under occlusion, eye-relief change, and pupil-size variation, tissue-linked polarization features in the sclera and cornea remain available when pupil contours and glints are weakened. The larger gains in the higher temporal camera placement are presented as support for this interpretation because that placement sees less of the eye opening and experiences more eyelid or eyelash occlusion. The head-mounted work extends the argument by emphasizing that its camera sees the top part of the eye, a view especially prone to eyelid and eyelash occlusion, and reports consistent PET advantages across baseline, linear calibration, Siamese personalization, and combined Siamese-plus-linear-calibration protocols [2511.04652; 2603.25889].

The practical significance claimed in the benchtop study is that PET may offer a simple, robust sensing modality for AI glasses, AR/VR/XR devices, compact always-on wearable eye trackers, and robust eye-based input for human-computer interaction. What is explicitly demonstrated is a benchtop non-form-factor station with single-camera, single-illuminator sensing per eye and matched-capacity improvements over intensity-only baselines. What remains an implication rather than a demonstrated result is fully integrated wearable deployment, including cost, power, and production-grade robustness [2511.04652].

Several limitations are explicitly identified. The benchtop paper notes a PFA resolution trade-off, because polarization mosaics trade spatial and angular sampling and demosaicking can reduce per-channel SNR. It also notes angular dependence of polarization efficiency, with measured AoLP and DoLP depending on incidence angle and sensor–illuminator orientation, and it emphasizes that compact wearable deployment will require management of stray polarization, optical coatings, stray light suppression, and mechanical tolerances. Only one NIR wavelength, \(850 \text{ nm}\), and one linear polarization illumination state are used. Quantitative breakdowns are not provided for eyeglasses, contact lenses, demographic subgroups, head pose variation, uncontrolled ambient lighting, or pathology and generalization effects [2511.04652].

The personalization paper identifies complementary limitations. It is a benchmarking study on one dataset and one imaging geometry; it does not deeply analyze hardware tradeoffs such as power, cost, exposure constraints, or broader deployment conditions; it still requires some calibration anchors even though the burden is reduced; and naive Siamese inference requires one forward pass per anchor, implying about \(9\times\) the baseline cost at 9 anchors before caching. It also does not include feature-visualization or interpretability experiments proving exactly which polarization structures drive the gain [2603.25889].

Future directions in the benchtop study include miniaturized polarimetric sensors with low interpixel crosstalk and high throughput or fill factor, alternative optical implementations such as metasurface routers or splitters, richer illumination schemes including temporal multiplexing of linear states and possibly circular polarization, algorithms that model personalization more explicitly, self-supervised learning using relationships among intensity, AoLP, and DoLP, and possible use of birefringence-linked ocular contrast for health monitoring. The head-mounted study suggests smarter anchor selection, further few-shot or meta-learning approaches, and deployment strategies such as caching anchor features or reducing anchor count to 3–5 when latency is critical. Taken together, these proposals situate PET as a technical program at the intersection of polarimetric imaging, near-eye sensing, and personalized gaze estimation rather than as a finalized device architecture [2511.04652; 2603.25889].

Source: https://www.emergentmind.com/topics/polarization-enabled-eye-tracking-pet