---
title: Real-Time Confocal Tracking
url: https://www.emergentmind.com/topics/real-time-confocal-tracking
type: topic
---

# Real-Time Confocal Tracking

Searching arXiv for recent and foundational papers on real-time confocal tracking.
Real-time confocal tracking denotes a family of confocal microscopy methods in which photon acquisition, position estimation, and mechanical or optical recentering are executed in closed loop, so that an emitter, tracer, or nanoparticle is kept within the effective detection volume while measurements continue. In the recent literature, the term spans several closely related regimes: feedback-driven single-particle tracking with single-pixel detection, high-speed single-beam rescanned confocal imaging synchronized with three-dimensional particle tracking, image-free 3D focus locking on point-like emitters, and asynchronous image-scanning implementations adapted to particle tracking [2508.13668]. Across these variants, the core objective is not merely imaging, but maintaining spatial registration under diffusion, drift, or rapid sample dynamics while preserving the background rejection and axial sectioning characteristic of confocal detection.

## 1. Conceptual scope and optical foundations

Confocal tracking systems share a common optical backbone: a tightly focused excitation beam is delivered through a high-NA objective, emitted or reflected light is separated from the excitation by a dichroic element, and out-of-focus background is rejected by a confocal pinhole before detection. The 2025 perspective describes this in generic form as a single-mode laser with $\lambda_{\mathrm{ex}} \simeq 488$–$640\ \mathrm{nm}$ coupled into the back aperture of a high-NA objective ($NA = 1.2$–$1.45$), with fluorescence routed through a confocal pinhole of $\sim 1$–$1.5$ Airy units to one or more APD or SPAD detectors [2508.13668]. Kim et al. implemented this architecture for nitrogen-vacancy centers using a 532 nm continuous-wave laser, high-NA oil-immersion objective, single-mode fiber confocal pinhole, and a 50:50 single-mode fiber beam splitter feeding two APDs in a Hanbury-Brown–Twiss arrangement [1801.02619].

The scanning and detection layer differentiates subfamilies of the method. In deterministic real-time feedback-driven tracking, the focal spot follows a predefined pattern such as a raster, Lissajous scan, or orbit around the nominal center; in constellation scanning, as in MINFLUX, the beam dwells at a small set of points arranged around the center [2508.13668]. By contrast, the high-speed rescanned confocal microscope of Klaassen et al. uses a single 532 nm laser beam, a “probe-scan” unit at the sample, and a second “re-scan” unit after the confocal pinhole that projects the image directly onto a 1024×1024 camera without ever splitting the beam in the sample [1912.06543]. Ta et al. modified a label-free image scanning microscope so that a resonant mirror oscillating at 12 kHz and a chromatic line generate asynchronous two-dimensional imaging at 24 kHz while retaining the lateral resolution gain and background rejection of regular label-free ISM [2308.15817].

A notable feature of the field is that “real time” does not imply a single sensing modality. Some systems infer position from temporally modulated single-photon counts, some reconstruct images on a camera, and some operate in an imaging-free manner from a few scalar intensity samples. The cited literature therefore treats real-time confocal tracking as a control-and-estimation problem built on confocal signal formation rather than as one fixed instrument topology.

## 2. Feedback loops, estimators, and control laws

The canonical real-time loop is scan $\rightarrow$ measure $\rightarrow$ estimate $\rightarrow$ move. In orbital tracking, the beam position is prescribed as
$$
x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad
y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),
$$
so an offset particle generates a sinusoidally modulated signal whose amplitude and phase encode $\Delta x$ and $\Delta y$ [2207.02210]. In the Knight’s Tour method, the excitation spot visits a discrete grid of points $r_{bi}$ and the estimated position is
$$
r_{\mathrm{est}}=\sum_{i=1}^{K} p_i \cdot r_{bi}, \qquad p_i = n_i/\sum_j n_j,
$$
with photon counts assigned to the scan points in sequence [2207.02210].

Kim et al. implemented a lock-in-based feedback controller around the local intensity gradient. Around the focus maximum $(x_0,y_0,z_0)$, the detected signal is expanded as
$$
I(t) \simeq I_0 + \left(\frac{\partial I}{\partial x}\right)A_x\cos(\Omega_x t)
+ \left(\frac{\partial I}{\partial y}\right)A_y\sin(\Omega_x t)
+ \left(\frac{\partial I}{\partial z}\right)A_z\cos(\Omega_z t),
$$
with typical modulation amplitudes $A_x=A_y=A_z=10\ \mathrm{nm}$, $\Omega_x/2\pi \approx 25\ \mathrm{Hz}$, and $\Omega_z/2\pi \approx 35\ \mathrm{Hz}$. The NanoTrak lock-in generates error signals proportional to the spatial derivatives, $E_x \propto (\partial I/\partial x)A_x$, $E_y \propto (\partial I/\partial y)A_y$, and $E_z \propto (\partial I/\partial z)A_z$, which are multiplied by a negative feedback gain set to $G=100$ and applied to the piezo stage [1801.02619].

The same paper formulates the controller in Laplace form as
$$
G(s)=K_p+\frac{K_i}{s}+K_d s,
$$
while noting that the implemented NanoTrak behaves essentially as a proportional–integral controller with fixed overall gain $\sim 100$ [1801.02619]. The broader comparative literature treats PID and Kalman filtering as the principal control classes. The perspective article gives the standard PID law
$$
u(t)=K_p\,e(t)+K_i\!\int_0^t e(t')\,dt' + K_d\,\frac{de}{dt},
$$
and also summarizes an Extended Kalman filter with state vector $x_k=[x,y,v_x,v_y]^T$, process covariance $Q$, and measurement covariance $R$ [2508.13668].

FiND represents a different control philosophy. Rather than demodulating an imposed periodic scan, it samples the confocal signal at six points around the current estimate and computes a finite-difference update,
$$
\mathbf D_k=\lambda\sum_{j\in\{x,y,z\}} \Bigl[s(\mathbf r_k+\delta\,\mathbf e_j)-s(\mathbf r_k-\delta\,\mathbf e_j)\Bigr]\mathbf e_j,
$$
followed by
$$
\mathbf r_{k+1}=\mathbf r_k+\mathbf D_k.
$$
The method uses no hardware add-ons beyond a standard confocal microscope and is explicitly presented as “imaging-free” and “non-trained” [2311.06479].

## 3. High-speed rescanning and synchronized three-dimensional tracking

A distinctive branch of real-time confocal tracking couples confocal image formation directly to dynamic readout. Klaassen et al. built a double-scan architecture whose heart is a single 532 nm laser beam scanned through the sample by a probe-scan unit and then re-scanned onto a two-dimensional camera after a 20 $\mu$m confocal pinhole. The probe and re-scan units each use a resonant mirror in $x$ and a galvanometric mirror in $y$, phase-locked in frequency and phase, with independently adjustable amplitudes [1912.06543]. Because only one beam is ever in the sample, the paper states that no inter-beam cross-talk occurs and that image fidelity is independent of frame rate [1912.06543].

In this system, the resonant mirrors oscillate at $f_r=16\ \mathrm{kHz}$, each line is scanned twice per period, and the effective line rate is therefore $f_{\mathrm{line}} \simeq 32\ \mathrm{kHz}$. With 1024 pixels in $x$, the pixel dwell time is
$$
\tau_{\mathrm{px}}=\frac{1}{f_{\mathrm{line}}\cdot 1024}\simeq 30\ \mathrm{ns},
$$
and scanning a diffraction-limited spot of approximately 4 pixels gives $\tau_{\mathrm{dwell}} \simeq 122\ \mathrm{ns}$ [1912.06543]. The frame rate follows
$$
f_{\mathrm{frame}}=\frac{f_{\mathrm{line}}}{N_y},
$$
so reducing the number of scanned lines $N_y$ increases the frame rate from video rates to approximately $1000\ \mathrm{Hz}$ for $N_y=32$ [1912.06543].

The same instrument can be extended by astigmatism particle tracking velocimetry (APTV). The fluorescence is split off by a dichroic mirror just after the objective, filtered, and directed through a cylindrical lens of focal length $f=50\ \mathrm{mm}$ onto a second camera, producing astigmatic particle images whose elliptical widths are calibrated against axial position:
$$
w_x(z)=w_0\sqrt{1+\bigl((z-z_0)/z_{Rx}\bigr)^2}, \qquad
w_y(z)=w_0\sqrt{1+\bigl((z-z_0)/z_{Ry}\bigr)^2}.
$$
With $\Delta w = w_x-w_y$, the system numerically or polynomially inverts to obtain $z=f(\Delta w)$, and near the midplane uses the linearized form $z \simeq \alpha\cdot\Delta w + \beta$ [1912.06543]. Both the confocal camera and the APTV camera are triggered from the same clock that drives the resonant and galvo scanners, and a single software pipeline reconstructs high-resolution confocal images while simultaneously extracting $x,y$ by centroiding and $z$ from $\Delta w$ in real time for frame rates up to several hundred Hz [1912.06543].

Ta et al. realized a related but label-free, camera-based direction. Their asynchronous ISM architecture uses a 12 kHz resonant mirror, chromatic scanning in the orthogonal direction, and optical photon reassignment. The emission path is rescanned with double amplitude, so one full super-resolution image is formed per resonant half-period, yielding a 24 kHz frame rate without electronic synchronization [2308.15817]. In this sense, real-time confocal tracking is not restricted to point-detector feedback loops; it also includes high-speed optically reassigned camera systems that support single-particle tracking.

## 4. Quantitative performance regimes

The literature reports performance in terms of localization precision, frame or update rate, tracking range, diffusion ceiling, dwell time, and photostability. The resulting parameter space is heterogeneous because the instruments are optimized for different tasks: diffraction-limited structural imaging, nanometric focus locking, fast diffusion tracking, or label-free nanoparticle tracking.

| Modality | Operating regime | Reported metrics |
|---|---|---|
| Single-beam re-scan confocal + APTV | $f_{\mathrm{line}}\simeq 32\ \mathrm{kHz}$; $N_y=32 \Rightarrow f_{\mathrm{frame}}=1000\ \mathrm{Hz}$ | lateral $\simeq 0.235\ \mu\mathrm{m}$; axial $\simeq 0.57\ \mu\mathrm{m}$; APTV precision $\sim 50$–$100\ \mathrm{nm}$ laterally and $\sim 100$–$200\ \mathrm{nm}$ axially over a $\pm 50\ \mu\mathrm{m}$ range [1912.06543] |
| NV-center lock-in tracking | update every $\approx 3\ \mathrm{ms}$; feedback bandwidth $f_c \approx 10$–$20\ \mathrm{Hz}$ | upper bound recovery time of $0.9\ \mathrm{s}$ upon $250\ \mathrm{nm}$ step shift; tracking amplitude $10\ \mathrm{nm}$; photon counts held within $\pm 5\%$ rms over $10$–$30\ \mathrm{h}$ [1801.02619] |
| FiND image-free tracking | tens to hundreds of Hz | $\Delta x \approx 9\ \mathrm{nm}$, $\Delta y \approx 8\ \mathrm{nm}$, $\Delta z \approx 9\ \mathrm{nm}$; convergence in a few iterations ($\le 10$) for $\mathrm{SNR} \ge 5$ [2311.06479] |
| Label-free asynchronous ISM | 24 kHz imaging; 1 kHz SPT acquisition | $FWHM_{\mathrm{conf}} = 240 \pm 20\ \mathrm{nm}$; $FWHM_{\mathrm{ISM}} = 170 \pm 10\ \mathrm{nm}$; $\sigma_{\mathrm{loc}} \simeq 4.3\ \mathrm{nm}$ for $N_p=3500$ photons [2308.15817] |
| RT-FD-SPT comparison | 8.33 kHz scan; $\Gamma_0 = 12.5\ \mathrm{kcps}$; $\mathrm{SBR}=10$ | orbital $D_{\max}\approx 2.5\ \mu\mathrm{m}^2/\mathrm{s}$; Knight’s Tour $D_{\max}\approx 20\ \mu\mathrm{m}^2/\mathrm{s}$; MINFLUX $D_{\max}\approx 0.2\ \mu\mathrm{m}^2/\mathrm{s}$ [2207.02210] |

The image-forming rescanned confocal microscope reports a theoretical confocal lateral resolution
$$
r_{\mathrm{lat}} \simeq 0.51\,\lambda/NA,
$$
which with $\lambda = 532\ \mathrm{nm}$ and $NA = 1.35$ gives $r_{\mathrm{lat,theo}} \simeq 0.20\ \mu\mathrm{m}$, while the experimental value inferred from the re-scanned spot size is $r_{\mathrm{lat}} \simeq 0.235\ \mu\mathrm{m}$. The corresponding axial expression
$$
r_{\mathrm{ax}} \simeq \frac{0.88\,\lambda}{n-\sqrt{n^2-NA^2}}
$$
with oil refractive index $n=1.518$ yields $r_{\mathrm{ax,theo}} \simeq 0.57\ \mu\mathrm{m}$ [1912.06543]. Under dense fluorescent colloid conditions, however, the same paper reports degraded effective resolution of lateral $\simeq 1.3\ \mu\mathrm{m}$ and axial $\simeq 4.1\ \mu\mathrm{m}$ because of reduced signal-to-noise and imperfect bead fluorescence [1912.06543].

The feedback-tracking literature makes the precision–speed compromise explicit. The 2022 theoretical comparison states that there is a fundamental trade-off between precision and speed: the Knight’s Tour method can track the fastest diffusion but with low precision, whereas MINFLUX is the most precise but only tracks slow diffusion [2207.02210]. The 2025 perspective gives representative update-rate bands of $f_{\mathrm{update}}\sim 1$–$5\ \mathrm{kHz}$ for deterministic or raster scanning with $\tau \approx 50$–$200\ \mu\mathrm{s}$, and $f_{\mathrm{update}}\approx 5$–$15\ \mathrm{kHz}$ for 6-point MINFLUX with $\tau \approx 10$–$30\ \mu\mathrm{s}$ [2508.13668]. This comparative framework is central to interpreting “real-time” claims, because temporal bandwidth, not just static localization precision, determines whether a particle remains trackable.

## 5. Signal, noise, and photophysics

Photon statistics and background rejection determine both localization precision and closed-loop stability. For the lock-in NV-center system, the photon-shot-noise-limited detectable displacement is written as
$$
\Delta x_{\min} \simeq \frac{\sqrt{I_0}}{\partial I/\partial x},
$$
and for $I_0 \approx 10^3/\mathrm{s}$ and $\partial I/\partial x \approx 0.1$ counts/ms per nm, the paper obtains $\Delta x_{\min} \approx 10\ \mathrm{nm}$, consistent with the chosen tracking amplitude [1801.02619]. The same work defines
$$
\mathrm{SNR}=\frac{I_p-I_b}{\sigma},
$$
with $I_p$ the mean photon count at focus, $I_b$ the background, and $\sigma$ the standard deviation in 1 ms; for a single NV center with $I_p\approx 5\times 10^3\ \mathrm{s}^{-1}$ and $I_b\approx 1$ count/ms, it reports a lowest SNR of 2.2 [1801.02619].

The theoretical and perspective papers phrase the same issue more generally. For two-point differential detection in one dimension, the perspective states a Cramér–Rao scaling
$$
\sigma_x \gtrsim \frac{R}{\sqrt{N_{\mathrm{tot}}}},
$$
where $R$ is the constellation radius and $N_{\mathrm{tot}}$ is the total photon count per update [2508.13668]. The 2022 comparison gives an approximate dynamic tracking error
$$
e^2 \approx D\tau + \sigma_{\mathrm{int}}^2,
$$
with transition to free-diffusion behavior $e \approx \sqrt{2D\tau}$ when the particle outruns the scan; it also states a cutoff diffusion $D_{\max}\sim (fR)^2/(2\pi^2\eta)$ above which tracking fails [2207.02210]. These results formalize the dependence of closed-loop performance on photon rate $\eta$, scan radius $R$, and feedback interval $\tau$.

Photophysics enters in different ways across implementations. The single-beam rescanned confocal microscope emphasizes ultrashort illumination per spot: at $\tau_{\mathrm{dwell}} \approx 122\ \mathrm{ns}$ and focal-spot intensity near $11\ \mathrm{mW}/\mu\mathrm{m}^2$, fluorescent beads could be imaged more than 8000 times before noticeable bleaching, versus a few hundred scans for systems with microsecond-level dwell times [1912.06543]. By contrast, the single-molecule tracking perspective lists higher instantaneous laser intensity and repeated excitation as a limitation of real-time confocal tracking, because they raise photobleaching and photodamage risks [2508.13668]. The apparent difference reflects the fact that photobleaching depends on the full excitation protocol, especially dwell time, peak intensity, and whether the system concentrates photons into sparse high-information scan positions or sweeps continuously over an image field.

## 6. Applications, limitations, and prospective directions

Applications fall into at least three broad groups. First are soft-matter and flow problems requiring simultaneous structure and kinematics. The single-beam + re-scan + APTV architecture is presented as ideal for dense colloidal suspensions under shear, microfluidic flows and mixing, dynamic wetting and dewetting at three-phase contact lines, and fast cellular processes such as organelle transport [1912.06543]. Second are point-emitter and quantum-defect problems, where the aim is long-term stabilization under low photon budgets. Kim et al. demonstrate long-term position tracking compatible with pulsed measurements of NV centers, including single spin magnetic resonance and Rabi oscillations, and explicitly note use-cases in NV optical trapping, NV tracking in fluid dynamics, and biological sensing using NV centers inside a biological cell [1801.02619]. Third are nanoparticle and label-free tracking regimes. Ta et al. track gold particles down to 20 nm and silica particles down to 50 nm, and image freely moving *Lactobacillus* with improved resolution [2308.15817].

Several limitations recur. In single-molecule RT-FD-SPT, only one molecule can be tracked at a time, so crowded samples require prior dilution or activation [2508.13668]. Hardware complexity is also substantial in many systems, especially when AODs, EODs, FPGA or lock-in electronics, fast stages, and low-latency computing are needed [2508.13668]. The single-beam rescanned confocal microscope notes that the objective back aperture is not fully filled, producing slight resolution loss; that z-scan speed is limited by the piezo and oil coupling; and that the 14 $\mu$m camera pixel size is somewhat larger than the re-scanned spot [1912.06543]. The label-free ISM system reports only $\sim 20\%$ throughput from sample to sensor and lacks active high-speed 3D scanning, so axial extension would require electrically tunable lenses or deformable mirrors [2308.15817]. FiND, while avoiding hardware add-ons, remains limited in loop rate by photon-count integration time and piezo settling, and in the reported implementation runs at tens to hundreds of Hz [2311.06479].

Prospective directions are correspondingly diverse. For rescanned confocal imaging, suggested upgrades include a liquid or electrowetting lens for kHz-level z-scanning, cooled higher-QE cameras with pixel sizes matched to the re-scanned spot, multicolor fluorescence via additional dichroics and beam paths, and incorporation of two-photon excitation [1912.06543]. For lock-in NV tracking, proposed optimizations include replacing the SR400 with an FPGA-based counter and fast D/A converter to increase bandwidth toward $100\ \mathrm{Hz}$, adding a derivative term to reduce overshoot, and using an acousto-optic deflector instead of a piezo stage for faster response [1801.02619]. The comparative RT-FD-SPT literature points toward Fisher-information-driven control, adaptive scan pattern selection, multimodal fluorescence plus iSCAT, and extension to true 3D through axial beam modulation or multi-plane detection [2207.02210]. The 2025 perspective additionally identifies parallelization and artificial intelligence as emerging efforts aimed at higher spatiotemporal resolution and greater computational and data efficiency [2508.13668].

Taken together, these studies show that real-time confocal tracking is not a single protocol but a technically diverse class of confocal feedback systems. Its unifying principle is the continuous conversion of sparse or image-based confocal measurements into low-latency spatial corrections. The principal design variables—scan geometry, dwell time, detector modality, estimator, actuator bandwidth, and photophysical budget—set the achievable compromise among localization precision, temporal resolution, axial range, and measurement duration.

Source: https://www.emergentmind.com/topics/real-time-confocal-tracking