---
title: Physical Frames Per Second (PhyFPS)
url: https://www.emergentmind.com/topics/physical-frames-per-second-phyfps
type: topic
---

# Physical Frames Per Second (PhyFPS)

Physical Frames Per Second (PhyFPS) denotes a physically meaningful temporal sampling or update rate: the rate at which distinct states are actually acquired, delivered, or implied by motion, rather than a nominal playback rate, container header, or algorithmically reconstructed cadence. In the literature, the term separates true sensor acquisition from temporal binning or sliding-window reuse, distinguishes a video’s intrinsic time base from its metadata FPS, and distinguishes physical display cadence from lower-rate model updates or repeated presents [2409.17180] [2603.14375] [1705.10930] [2502.20934]. This suggests that PhyFPS is not merely a camera specification, but a general temporal quantity spanning measurement systems, video models, and perception–action pipelines.

## 1. Definition and conceptual scope

A foundational formulation appears in ultrafast imaging, where temporal resolution per frame is denoted by $\Delta t$ and the physical sampling rate is written as
$$
\mathrm{fps} \equiv \mathrm{PhyFPS} = \frac{1}{\Delta t}.
$$
In that usage, PhyFPS is explicitly not the playback rate; it is the rate at which the physical event is sampled during acquisition [1808.00428].

Subsequent work generalizes the distinction. In visual chronometry, PhyFPS is defined as the intrinsic frame rate implied by motion in a video, in contrast to nominal display or encoding FPS recorded in container metadata. That literature separates $f_{\mathrm{capture}}$, $f_{\mathrm{meta}}$, $f_{\mathrm{playback}}$, and $f_{\mathrm{phy}}$, and argues that meta FPS is often a convention while PhyFPS is determined by the motion itself [2603.14375]. In real-time surgical segmentation, the distinction is operationalized differently: the physical stream remains at 25 FPS, while the model’s processing FPS $f$ may be lower, so predictions are updated every $\Delta_f = 25/f$ frames and held constant in between [2502.20934].

Several systems explicitly contrast physical acquisition with computationally inflated rates. In retinal Doppler holography, interferograms are physically recorded at 33 kHz by the sensor, with “no temporal binning or computational frame-rate reconstruction” [2409.17180]. In $\mu$FTP, the system produces “pseudo” 20,000 3D fps via sliding-window reconstruction, but the physically independent PhyFPS is 10,000 because two projected patterns are required per independent 3D frame [1705.10930]. In competitive rendering, a further distinction appears between engine FPS, display Hz, and the effective rate at which distinct, actionable game states actually reach the viewer; this rate is formalized as a physical update rate rather than a raw render count [2208.11774].

## 2. Temporal sampling, Nyquist limits, and motion blur

The physical significance of PhyFPS is inseparable from temporal resolution, bandwidth, and blur. In light-in-flight photography, the motion-freeze criterion is stated as
$$
\Delta t < \frac{p}{v}
\quad \Rightarrow \quad
\mathrm{fps} > \frac{v}{p},
$$
where $v$ is object speed and $p$ is the projected size of one pixel on the object. The same review states a sampling-theorem condition,
$$
\mathrm{fps} \equiv \frac{1}{\Delta t} \ge 2 f_{\max},
$$
for a process with characteristic temporal bandwidth $f_{\max}$ [1808.00428].

Retinal Doppler holography provides a concrete biomedical instance of these constraints. With PhyFPS $=33{,}000\ \mathrm{s}^{-1}$, the per-frame exposure is approximately $\Delta t \approx 30.3\ \mu\mathrm{s}$, which suppresses motion blur at the fundus and preserves high-frequency Doppler content. Uniform sampling at $33{,}000\ \mathrm{s}^{-1}$ sets a Nyquist-limited Doppler bandwidth of $f_{\mathrm{Nyquist}} \approx 16.5\ \mathrm{kHz}$. Using 512-frame STFT windows yields a time span per spectrum of approximately $15.5\ \mathrm{ms}$ and a frequency bin spacing of approximately $64.5\ \mathrm{Hz}$. Under the forward scattering model,
$$
v = \lambda \Delta f / N,
$$
with $\lambda = 852\ \mathrm{nm}$ and $N \approx 0.124$, the Nyquist limit implies $v_{\max} \approx 113\ \mathrm{mm/s}$ and the STFT bin width implies a velocity resolution of approximately $0.44\ \mathrm{mm/s}$ [2409.17180].

Rolling-shutter compressive imaging makes the temporal sampling geometry explicit at the sensor-row level. For an Andor Zyla 4.2 sCMOS camera operated in fast line readout mode, the measured line time is $\Delta t_{\mathrm{row}} \approx 9.6\ \mu\mathrm{s}$, implying
$$
\mathrm{PhyFPS} = \frac{1}{\Delta t_{\mathrm{row}}} \approx 104{,}166\ \mathrm{s}^{-1}.
$$
With a $108 \times 108$ ROI and dual rolling shutter, one captured frame contains $54$ distinct physical temporal samples over a total time window of approximately $0.5184\ \mathrm{ms}$ [2004.09614].

In event-based structured light, the limiting mechanism is not the light scanner but the event throughput. For full-frame swept-plane depth capture, the bandwidth relation is
$$
F_{\mathrm{full}} \le \frac{R_e}{E_{\mathrm{frame}}},
$$
and with $R_e \approx 10^6$ events/s and $E_{\mathrm{frame}} \approx 720$, the theoretical full-frame limit is about $1.39 \times 10^3\ \mathrm{Hz}$, with practical operation near $1000\ \mathrm{fps}$. In ROI scanning, reducing the active fraction of rows increases PhyFPS approximately as $F_{\mathrm{ROI}} = F_{\mathrm{full}}/f$ [2411.18597].

## 3. Biomedical and 3D sensing implementations

In retinal hemodynamics, PhyFPS functions as the physical determinant of Doppler bandwidth, exposure time, and physiological observability. The reported system uses an Ametek Phantom S711 streaming camera at 33 kHz with $384 \times 384$ pixels and $20\ \mu\mathrm{m}$ pixel pitch, producing a 16-bit interferogram stream of the eye fundus. Real-time rendering is achieved by Fresnel transform and PCA on stacks of 16 consecutive holograms using Holovibes, while raw interferograms are saved concurrently via Euresys Coaxlink QSFP+. In one acquisition, $131{,}072$ frames were recorded over approximately $3.8$–$4.0\ \mathrm{s}$, yielding $18\ \mathrm{GB}$ of data and a sustained throughput on the order of $4.7\ \mathrm{GB/s}$ [2409.17180].

The quantitative pipeline in that work couples physical sampling to blood-flow inference. Primary in-plane retinal arteries are segmented, local Doppler broadening is computed relative to surrounding tissue in the $6$–$16.5\ \mathrm{kHz}$ band, and the forward-scatter relation $v = \lambda \Delta f / N$ is used to estimate local RMS blood velocity. Volumetric flow is then computed as
$$
Q = \int_A v(x,y)\, dA,
\qquad
Q \approx \sum_i v_i A_i.
$$
In the reported control subject, the mean total retinal arterial blood volume rate is $35\ \mu\mathrm{L/min}$, a value described as commensurate with bidirectional laser Doppler velocimetry and Doppler FD-OCT [2409.17180].

In high-speed structured-light 3D sensing, $\mu$FTP defines physical 3D frame rate by the number of projected patterns required per independent reconstruction. With projector and camera both operated at $20{,}000\ \mathrm{fps}$ and $N=2$ projected patterns per 3D frame, the paper states
$$
\mathrm{PhyFPS} = \min(R_p,R_c)/N = 20{,}000/2 = 10{,}000\ \mathrm{fps}.
$$
The pattern switching period is approximately $50\ \mu\mathrm{s}$, the single phase-carrying exposure is $46\ \mu\mathrm{s}$, and one independent 3D reconstruction is produced every $100\ \mu\mathrm{s}$. Because the phase is encoded in a single image, the method is described as motion-artifact-free and explicitly distinguishes true 10,000 3D fps from pseudo 20,000 fps obtained by reusing overlapping frames [1705.10930].

Event-camera structured light extends the same principle into a different hardware regime. The acousto-optic scanner reaches up to $2 \times 10^6$ light planes/s, but full-frame depth capture remains near $1000\ \mathrm{fps}$ because the event camera’s full-frame bandwidth is the bottleneck. Since one sweep of a single light plane produces one full-frame depth map, the paper states that PhyFPS equals the sweep frequency, provided the camera and link can keep up. ROI-only scanning reduces events per sweep and permits approximately $10\ \mathrm{kHz}$ operation in the prototype [2411.18597].

## 4. Ultrafast optical imaging

In ultrafast optical imaging, PhyFPS enters the picosecond- to femtosecond-scale regime. A review of light-in-flight photography surveys techniques with temporal resolution from picoseconds to femtoseconds, corresponding to PhyFPS in the $10^{11}$–$10^{14}$ range. At $10^{12}\ \mathrm{fps}$, $\Delta t = 1\ \mathrm{ps}$ and light travels approximately $0.3\ \mathrm{mm}$ per frame; at $10\ \mathrm{ps}$, the distance is about $3\ \mathrm{mm}$ [1808.00428].

Framing integration photography (FIP) with an inversed 4f system presents one of the clearest statements of “physical” framing. The system generates multiple independently time-gated probe sub-pulses from a single incident femtosecond pulse by a stepped delay element and a lenslet array, and maps each delayed sub-pulse to a disjoint detector region without compressive inversion or deconvolution. With fused silica delay steps of thickness increment $\Delta h = 0.12\ \mathrm{mm}$ and group index $n_g = 1.4671$ at $800\ \mathrm{nm}$, the per-frame delay is
$$
\Delta t = (n_g - 1)\Delta h / c = 187\ \mathrm{fs},
$$
yielding
$$
f_{\mathrm{phy}} = 1/\Delta t = 5.3 \times 10^{12}\ \mathrm{fps}.
$$
Four frames were captured at nominal times $0$, $187$, $374$, and $561\ \mathrm{fs}$, and the measured intrinsic object-space spatial resolution was $110.4\ \mathrm{lp/mm}$ [2110.01941].

Multiple non-collinear optical parametric amplifiers (MOPA) achieve a comparable but distinct form of physical framing. Four OPA channels, each with its own femtosecond pump delay, generate four physically distinct frames in one shot at equal 100 fs spacing, giving
$$
\mathrm{FPS} = 1/100\ \mathrm{fs} = 10^{13}\ \mathrm{fps}.
$$
The exposure time is approximately $40\ \mathrm{fs}$, the frames are recorded on four separate CCDs, and the spatial resolution exceeds $30\ \mathrm{lp/mm}$, with visible features up to $36\ \mathrm{lp/mm}$ in the high-resolution configuration [1807.00685].

These ultrafast systems also sharpen the distinction between physical and reconstructed frame rates. FIP explicitly contrasts itself with CUP and streak-camera scans because its time-separated information is optically decoded to disjoint detector regions without computational temporal demixing [2110.01941]. The MOPA work likewise emphasizes that its frames are discrete, optically formed, and physically captured in a single shot, rather than reconstructed from a compressive measurement [1807.00685]. A plausible implication is that, in ultrafast imaging, the designation “physical” is tied not only to $\Delta t$ but also to how temporal separation is realized in hardware.

## 5. Video AI and algorithmic systems

In real-time video algorithms, PhyFPS often defines the physical cadence against which lower-rate computation must be interpreted. In zero-shot surgical video segmentation, the videos are physically captured at 25 FPS, so PhyFPS is fixed at 25 for Cholec80/CholecSeg8k. The model’s processing FPS is varied over $f \in \{1, 10, 15, 20, 25\}$, with an update-hold mechanism in which a prediction on frame $X_t$ persists for frames $X_{t+1}, \ldots, X_{t+\Delta_f-1}$ where $\Delta_f = 25/f$. Under sampled-frames evaluation, 1 FPS can slightly outperform 25 FPS because fewer frames smooth out segmentation inconsistencies, but under real-time streaming evaluation on all physical frames, higher processing FPS consistently improves IoU and temporal coherence for dynamic targets such as the grasper [2502.20934].

That same study turns PhyFPS into an evaluation principle. It shows that anchor-frame evaluation largely removes the apparent superiority of low FPS, and that professional respondents consistently prefer higher-FPS overlays; low-FPS overlays are described as “choppy” and “out-of-sync.” The paper therefore argues that systems intended for intraoperative display should be assessed on the full physical frame stream rather than on sparsely sampled subsets [2502.20934].

Visual Chronometer addresses a different but related problem: inferring a video’s physical time base from motion rather than metadata. The method predicts continuous-valued $f_{\mathrm{phy}}$ from visual dynamics, is trained by controlled temporal resampling, and distinguishes $f_{\mathrm{capture}}$, $f_{\mathrm{meta}}$, $f_{\mathrm{playback}}$, and $f_{\mathrm{phy}}$. On PhyFPS-Bench-Real, VC-Common reports Avg Pred $39.20$, MAE $3.46$, and MAPE $9\%$, while audits on PhyFPS-Bench-Gen report pervasive meta-vs-PhyFPS mismatch and nontrivial intra- and inter-video CVs in state-of-the-art generators [2603.14375].

High-frame-rate video understanding extends the same logic into multimodal LLMs. F-16 increases video understanding input rate to 16 FPS and compresses visual tokens within each 1-second clip. For a local window of width $w = 16$, per-frame tokens $Z_i$ are concatenated, passed through a high-frame-rate aligner, and post-pooled so that one 1-second window produces about $p/4$ tokens after post-pooling, independent of $F$. The paper reports Video-MME results of Avg $65.0$, Short $78.9$, Medium $63.2$, and Long $52.8$ for the 16 FPS system, and states that higher FPS yields large gains on motion-centric benchmarks such as TemporalBench [2503.13956].

Deployed machine vision systems add a systems interpretation of PhyFPS. In real-time detection, Physical Frames Per Second is defined as the number of camera frames the entire deployed system can accept, process, and produce detections for per second, including capture, pre-processing, inference, post-processing, and output. The reported detector exceeds 200 FPS on a GTX 1080 under its evaluation protocol, but the account explicitly notes that PhyFPS can degrade in deployment when camera I/O, decoding, or scheduling become bottlenecks [1805.06361].

## 6. Interactive pipelines, perception, and recurring misconceptions

In interactive systems, PhyFPS can be defined by the rate of complete perception–action cycles rather than the rate of image acquisition alone. FirstPersonScience proposes
$$
T_{\mathrm{cycle}} = T_{\mathrm{response}} + L_{\mathrm{total}},
\qquad
\mathrm{PhyFPS} = \frac{1}{T_{\mathrm{cycle}}} = \frac{1}{T_{\mathrm{response}} + L_{\mathrm{total}}},
$$
where $L_{\mathrm{total}}$ is click-to-photon latency and
$$
L_{\mathrm{added}} = N_{\mathrm{delay}} \cdot T_{\mathrm{frame}},
\qquad
T_{\mathrm{frame}} = \frac{1}{F_{\mathrm{display}}}.
$$
At 60 fps, injected delays of 0, 1, and 2 frames produce distinct latency modes separated by about $16.7\ \mathrm{ms}$, and in the training demonstration task completion time decreased from a mean of $1.78\ \mathrm{s}$ in the first 27 trials to $1.34\ \mathrm{s}$ in the final 28 trials while system latency was held constant [2202.06429].

Competitive rendering extends the same notion from laboratory tasks to esports pipelines. There, PhyFPS is the effective rate at which distinct, physically meaningful updates of game state are actually delivered to the viewer through input, simulation, rendering, presentation, scanout, and pixel response. The formal definition is
$$
R_{\mathrm{phy}} = \frac{1}{E[\Delta t_{\mathrm{update}}]},
$$
with a practical steady-state approximation
$$
R_{\mathrm{phy}} \approx p_{\mathrm{new}} \cdot \min\{H, F_{\mathrm{present}}, S\},
$$
where $H$ is display refresh, $F_{\mathrm{present}}$ is actual present rate, $S$ is simulation tick rate, and $p_{\mathrm{new}}$ is the fraction of refreshes that carry new simulation state [2208.11774].

Several recurring misconceptions follow directly from these definitions. One is that nominal or metadata FPS is equivalent to the physical time base of a video; the chronometric-hallucination literature rejects this by showing that slow-motion, normal-rate, and time-lapse footage may share the same meta FPS while encoding different $f_{\mathrm{phy}}$ [2603.14375]. A second is that low FPS can be judged adequate from sparse offline metrics; the surgical segmentation study shows that this conclusion can arise from sampled-frames bias and disappear under full streaming evaluation on the physical frame stream [2502.20934]. A third is that pseudo or reconstructed frame rates are interchangeable with physical acquisition rates; both retinal holography and $\mu$FTP state the opposite explicitly by distinguishing sensor-recorded cadence from computational reuse or temporal reconstruction [2409.17180] [1705.10930].

Taken together, these works define PhyFPS as a cross-domain measure of temporal fidelity. In acquisition systems it is set by exposure, gating, scanout, or sweep timing; in vision models it is the physical input cadence against which processing and evaluation must be defined; in interactive pipelines it is the rate at which new, decision-relevant state reaches a human observer. The shared theme is that PhyFPS is meaningful precisely because it is anchored to physical time rather than to nominal headers, display conventions, or algorithmic interpolation.

Source: https://www.emergentmind.com/topics/physical-frames-per-second-phyfps