Papers
Topics
Authors
Recent
Search
2000 character limit reached

Real-Time Confocal Tracking

Updated 9 July 2026
  • Real-Time Confocal Tracking is a microscopy method that uses focused excitation, confocal detection, and a closed-loop control system to keep emitters within the detection volume.
  • It employs diverse scanning strategies like orbital, Knight’s Tour, and FiND methods, each balancing precision, speed, and imaging resolution under varying conditions.
  • Advances in control algorithms, high-speed rescanning, and signal processing enable improved localization, robustness against drift, and adaptation to dense or dynamic samples.

Searching arXiv for recent and foundational papers on real-time confocal tracking. Real-time confocal tracking denotes a family of confocal microscopy methods in which photon acquisition, position estimation, and mechanical or optical recentering are executed in closed loop, so that an emitter, tracer, or nanoparticle is kept within the effective detection volume while measurements continue. In the recent literature, the term spans several closely related regimes: feedback-driven single-particle tracking with single-pixel detection, high-speed single-beam rescanned confocal imaging synchronized with three-dimensional particle tracking, image-free 3D focus locking on point-like emitters, and asynchronous image-scanning implementations adapted to particle tracking (Xu et al., 19 Aug 2025). Across these variants, the core objective is not merely imaging, but maintaining spatial registration under diffusion, drift, or rapid sample dynamics while preserving the background rejection and axial sectioning characteristic of confocal detection.

1. Conceptual scope and optical foundations

Confocal tracking systems share a common optical backbone: a tightly focused excitation beam is delivered through a high-NA objective, emitted or reflected light is separated from the excitation by a dichroic element, and out-of-focus background is rejected by a confocal pinhole before detection. The 2025 perspective describes this in generic form as a single-mode laser with λex488\lambda_{\mathrm{ex}} \simeq 488640 nm640\ \mathrm{nm} coupled into the back aperture of a high-NA objective (NA=1.2NA = 1.2–$1.45$), with fluorescence routed through a confocal pinhole of 1\sim 1–$1.5$ Airy units to one or more APD or SPAD detectors (Xu et al., 19 Aug 2025). Kim et al. implemented this architecture for nitrogen-vacancy centers using a 532 nm continuous-wave laser, high-NA oil-immersion objective, single-mode fiber confocal pinhole, and a 50:50 single-mode fiber beam splitter feeding two APDs in a Hanbury-Brown–Twiss arrangement (Kim et al., 2018).

The scanning and detection layer differentiates subfamilies of the method. In deterministic real-time feedback-driven tracking, the focal spot follows a predefined pattern such as a raster, Lissajous scan, or orbit around the nominal center; in constellation scanning, as in MINFLUX, the beam dwells at a small set of points arranged around the center (Xu et al., 19 Aug 2025). By contrast, the high-speed rescanned confocal microscope of Klaassen et al. uses a single 532 nm laser beam, a “probe-scan” unit at the sample, and a second “re-scan” unit after the confocal pinhole that projects the image directly onto a 1024×1024 camera without ever splitting the beam in the sample (Straub et al., 2019). Ta et al. modified a label-free image scanning microscope so that a resonant mirror oscillating at 12 kHz and a chromatic line generate asynchronous two-dimensional imaging at 24 kHz while retaining the lateral resolution gain and background rejection of regular label-free ISM (Ta et al., 2023).

A notable feature of the field is that “real time” does not imply a single sensing modality. Some systems infer position from temporally modulated single-photon counts, some reconstruct images on a camera, and some operate in an imaging-free manner from a few scalar intensity samples. The cited literature therefore treats real-time confocal tracking as a control-and-estimation problem built on confocal signal formation rather than as one fixed instrument topology.

2. Feedback loops, estimators, and control laws

The canonical real-time loop is scan \rightarrow measure \rightarrow estimate \rightarrow move. In orbital tracking, the beam position is prescribed as

xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),

so an offset particle generates a sinusoidally modulated signal whose amplitude and phase encode 640 nm640\ \mathrm{nm}0 and 640 nm640\ \mathrm{nm}1 (Heerden et al., 2022). In the Knight’s Tour method, the excitation spot visits a discrete grid of points 640 nm640\ \mathrm{nm}2 and the estimated position is

640 nm640\ \mathrm{nm}3

with photon counts assigned to the scan points in sequence (Heerden et al., 2022).

Kim et al. implemented a lock-in-based feedback controller around the local intensity gradient. Around the focus maximum 640 nm640\ \mathrm{nm}4, the detected signal is expanded as

640 nm640\ \mathrm{nm}5

with typical modulation amplitudes 640 nm640\ \mathrm{nm}6, 640 nm640\ \mathrm{nm}7, and 640 nm640\ \mathrm{nm}8. The NanoTrak lock-in generates error signals proportional to the spatial derivatives, 640 nm640\ \mathrm{nm}9, NA=1.2NA = 1.20, and NA=1.2NA = 1.21, which are multiplied by a negative feedback gain set to NA=1.2NA = 1.22 and applied to the piezo stage (Kim et al., 2018).

The same paper formulates the controller in Laplace form as

NA=1.2NA = 1.23

while noting that the implemented NanoTrak behaves essentially as a proportional–integral controller with fixed overall gain NA=1.2NA = 1.24 (Kim et al., 2018). The broader comparative literature treats PID and Kalman filtering as the principal control classes. The perspective article gives the standard PID law

NA=1.2NA = 1.25

and also summarizes an Extended Kalman filter with state vector NA=1.2NA = 1.26, process covariance NA=1.2NA = 1.27, and measurement covariance NA=1.2NA = 1.28 (Xu et al., 19 Aug 2025).

FiND represents a different control philosophy. Rather than demodulating an imposed periodic scan, it samples the confocal signal at six points around the current estimate and computes a finite-difference update,

NA=1.2NA = 1.29

followed by

$1.45$0

The method uses no hardware add-ons beyond a standard confocal microscope and is explicitly presented as “imaging-free” and “non-trained” (Sahoo et al., 2023).

3. High-speed rescanning and synchronized three-dimensional tracking

A distinctive branch of real-time confocal tracking couples confocal image formation directly to dynamic readout. Klaassen et al. built a double-scan architecture whose heart is a single 532 nm laser beam scanned through the sample by a probe-scan unit and then re-scanned onto a two-dimensional camera after a 20 $1.45$1m confocal pinhole. The probe and re-scan units each use a resonant mirror in $1.45$2 and a galvanometric mirror in $1.45$3, phase-locked in frequency and phase, with independently adjustable amplitudes (Straub et al., 2019). Because only one beam is ever in the sample, the paper states that no inter-beam cross-talk occurs and that image fidelity is independent of frame rate (Straub et al., 2019).

In this system, the resonant mirrors oscillate at $1.45$4, each line is scanned twice per period, and the effective line rate is therefore $1.45$5. With 1024 pixels in $1.45$6, the pixel dwell time is

$1.45$7

and scanning a diffraction-limited spot of approximately 4 pixels gives $1.45$8 (Straub et al., 2019). The frame rate follows

$1.45$9

so reducing the number of scanned lines 1\sim 10 increases the frame rate from video rates to approximately 1\sim 11 for 1\sim 12 (Straub et al., 2019).

The same instrument can be extended by astigmatism particle tracking velocimetry (APTV). The fluorescence is split off by a dichroic mirror just after the objective, filtered, and directed through a cylindrical lens of focal length 1\sim 13 onto a second camera, producing astigmatic particle images whose elliptical widths are calibrated against axial position:

1\sim 14

With 1\sim 15, the system numerically or polynomially inverts to obtain 1\sim 16, and near the midplane uses the linearized form 1\sim 17 (Straub et al., 2019). Both the confocal camera and the APTV camera are triggered from the same clock that drives the resonant and galvo scanners, and a single software pipeline reconstructs high-resolution confocal images while simultaneously extracting 1\sim 18 by centroiding and 1\sim 19 from $1.5$0 in real time for frame rates up to several hundred Hz (Straub et al., 2019).

Ta et al. realized a related but label-free, camera-based direction. Their asynchronous ISM architecture uses a 12 kHz resonant mirror, chromatic scanning in the orthogonal direction, and optical photon reassignment. The emission path is rescanned with double amplitude, so one full super-resolution image is formed per resonant half-period, yielding a 24 kHz frame rate without electronic synchronization (Ta et al., 2023). In this sense, real-time confocal tracking is not restricted to point-detector feedback loops; it also includes high-speed optically reassigned camera systems that support single-particle tracking.

4. Quantitative performance regimes

The literature reports performance in terms of localization precision, frame or update rate, tracking range, diffusion ceiling, dwell time, and photostability. The resulting parameter space is heterogeneous because the instruments are optimized for different tasks: diffraction-limited structural imaging, nanometric focus locking, fast diffusion tracking, or label-free nanoparticle tracking.

Modality Operating regime Reported metrics
Single-beam re-scan confocal + APTV $1.5$1; $1.5$2 lateral $1.5$3; axial $1.5$4; APTV precision $1.5$5–$1.5$6 laterally and $1.5$7–$1.5$8 axially over a $1.5$9 range (Straub et al., 2019)
NV-center lock-in tracking update every \rightarrow0; feedback bandwidth \rightarrow1–\rightarrow2 upper bound recovery time of \rightarrow3 upon \rightarrow4 step shift; tracking amplitude \rightarrow5; photon counts held within \rightarrow6 rms over \rightarrow7–\rightarrow8 (Kim et al., 2018)
FiND image-free tracking tens to hundreds of Hz \rightarrow9, \rightarrow0, \rightarrow1; convergence in a few iterations (\rightarrow2) for \rightarrow3 (Sahoo et al., 2023)
Label-free asynchronous ISM 24 kHz imaging; 1 kHz SPT acquisition \rightarrow4; \rightarrow5; \rightarrow6 for \rightarrow7 photons (Ta et al., 2023)
RT-FD-SPT comparison 8.33 kHz scan; \rightarrow8; \rightarrow9 orbital \rightarrow0; Knight’s Tour \rightarrow1; MINFLUX \rightarrow2 (Heerden et al., 2022)

The image-forming rescanned confocal microscope reports a theoretical confocal lateral resolution

\rightarrow3

which with \rightarrow4 and \rightarrow5 gives \rightarrow6, while the experimental value inferred from the re-scanned spot size is \rightarrow7. The corresponding axial expression

\rightarrow8

with oil refractive index \rightarrow9 yields xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),0 (Straub et al., 2019). Under dense fluorescent colloid conditions, however, the same paper reports degraded effective resolution of lateral xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),1 and axial xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),2 because of reduced signal-to-noise and imperfect bead fluorescence (Straub et al., 2019).

The feedback-tracking literature makes the precision–speed compromise explicit. The 2022 theoretical comparison states that there is a fundamental trade-off between precision and speed: the Knight’s Tour method can track the fastest diffusion but with low precision, whereas MINFLUX is the most precise but only tracks slow diffusion (Heerden et al., 2022). The 2025 perspective gives representative update-rate bands of xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),3–xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),4 for deterministic or raster scanning with xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),5–xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),6, and xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),7–xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),8 for 6-point MINFLUX with xbeam(t)=xest(t)+Rcos(2πft),ybeam(t)=yest(t)+Rsin(2πft),x_{\mathrm{beam}}(t)=x_{\mathrm{est}}(t)+R\cos(2\pi f t), \qquad y_{\mathrm{beam}}(t)=y_{\mathrm{est}}(t)+R\sin(2\pi f t),9–640 nm640\ \mathrm{nm}00 (Xu et al., 19 Aug 2025). This comparative framework is central to interpreting “real-time” claims, because temporal bandwidth, not just static localization precision, determines whether a particle remains trackable.

5. Signal, noise, and photophysics

Photon statistics and background rejection determine both localization precision and closed-loop stability. For the lock-in NV-center system, the photon-shot-noise-limited detectable displacement is written as

640 nm640\ \mathrm{nm}01

and for 640 nm640\ \mathrm{nm}02 and 640 nm640\ \mathrm{nm}03 counts/ms per nm, the paper obtains 640 nm640\ \mathrm{nm}04, consistent with the chosen tracking amplitude (Kim et al., 2018). The same work defines

640 nm640\ \mathrm{nm}05

with 640 nm640\ \mathrm{nm}06 the mean photon count at focus, 640 nm640\ \mathrm{nm}07 the background, and 640 nm640\ \mathrm{nm}08 the standard deviation in 1 ms; for a single NV center with 640 nm640\ \mathrm{nm}09 and 640 nm640\ \mathrm{nm}10 count/ms, it reports a lowest SNR of 2.2 (Kim et al., 2018).

The theoretical and perspective papers phrase the same issue more generally. For two-point differential detection in one dimension, the perspective states a Cramér–Rao scaling

640 nm640\ \mathrm{nm}11

where 640 nm640\ \mathrm{nm}12 is the constellation radius and 640 nm640\ \mathrm{nm}13 is the total photon count per update (Xu et al., 19 Aug 2025). The 2022 comparison gives an approximate dynamic tracking error

640 nm640\ \mathrm{nm}14

with transition to free-diffusion behavior 640 nm640\ \mathrm{nm}15 when the particle outruns the scan; it also states a cutoff diffusion 640 nm640\ \mathrm{nm}16 above which tracking fails (Heerden et al., 2022). These results formalize the dependence of closed-loop performance on photon rate 640 nm640\ \mathrm{nm}17, scan radius 640 nm640\ \mathrm{nm}18, and feedback interval 640 nm640\ \mathrm{nm}19.

Photophysics enters in different ways across implementations. The single-beam rescanned confocal microscope emphasizes ultrashort illumination per spot: at 640 nm640\ \mathrm{nm}20 and focal-spot intensity near 640 nm640\ \mathrm{nm}21, fluorescent beads could be imaged more than 8000 times before noticeable bleaching, versus a few hundred scans for systems with microsecond-level dwell times (Straub et al., 2019). By contrast, the single-molecule tracking perspective lists higher instantaneous laser intensity and repeated excitation as a limitation of real-time confocal tracking, because they raise photobleaching and photodamage risks (Xu et al., 19 Aug 2025). The apparent difference reflects the fact that photobleaching depends on the full excitation protocol, especially dwell time, peak intensity, and whether the system concentrates photons into sparse high-information scan positions or sweeps continuously over an image field.

6. Applications, limitations, and prospective directions

Applications fall into at least three broad groups. First are soft-matter and flow problems requiring simultaneous structure and kinematics. The single-beam + re-scan + APTV architecture is presented as ideal for dense colloidal suspensions under shear, microfluidic flows and mixing, dynamic wetting and dewetting at three-phase contact lines, and fast cellular processes such as organelle transport (Straub et al., 2019). Second are point-emitter and quantum-defect problems, where the aim is long-term stabilization under low photon budgets. Kim et al. demonstrate long-term position tracking compatible with pulsed measurements of NV centers, including single spin magnetic resonance and Rabi oscillations, and explicitly note use-cases in NV optical trapping, NV tracking in fluid dynamics, and biological sensing using NV centers inside a biological cell (Kim et al., 2018). Third are nanoparticle and label-free tracking regimes. Ta et al. track gold particles down to 20 nm and silica particles down to 50 nm, and image freely moving Lactobacillus with improved resolution (Ta et al., 2023).

Several limitations recur. In single-molecule RT-FD-SPT, only one molecule can be tracked at a time, so crowded samples require prior dilution or activation (Xu et al., 19 Aug 2025). Hardware complexity is also substantial in many systems, especially when AODs, EODs, FPGA or lock-in electronics, fast stages, and low-latency computing are needed (Xu et al., 19 Aug 2025). The single-beam rescanned confocal microscope notes that the objective back aperture is not fully filled, producing slight resolution loss; that z-scan speed is limited by the piezo and oil coupling; and that the 14 640 nm640\ \mathrm{nm}22m camera pixel size is somewhat larger than the re-scanned spot (Straub et al., 2019). The label-free ISM system reports only 640 nm640\ \mathrm{nm}23 throughput from sample to sensor and lacks active high-speed 3D scanning, so axial extension would require electrically tunable lenses or deformable mirrors (Ta et al., 2023). FiND, while avoiding hardware add-ons, remains limited in loop rate by photon-count integration time and piezo settling, and in the reported implementation runs at tens to hundreds of Hz (Sahoo et al., 2023).

Prospective directions are correspondingly diverse. For rescanned confocal imaging, suggested upgrades include a liquid or electrowetting lens for kHz-level z-scanning, cooled higher-QE cameras with pixel sizes matched to the re-scanned spot, multicolor fluorescence via additional dichroics and beam paths, and incorporation of two-photon excitation (Straub et al., 2019). For lock-in NV tracking, proposed optimizations include replacing the SR400 with an FPGA-based counter and fast D/A converter to increase bandwidth toward 640 nm640\ \mathrm{nm}24, adding a derivative term to reduce overshoot, and using an acousto-optic deflector instead of a piezo stage for faster response (Kim et al., 2018). The comparative RT-FD-SPT literature points toward Fisher-information-driven control, adaptive scan pattern selection, multimodal fluorescence plus iSCAT, and extension to true 3D through axial beam modulation or multi-plane detection (Heerden et al., 2022). The 2025 perspective additionally identifies parallelization and artificial intelligence as emerging efforts aimed at higher spatiotemporal resolution and greater computational and data efficiency (Xu et al., 19 Aug 2025).

Taken together, these studies show that real-time confocal tracking is not a single protocol but a technically diverse class of confocal feedback systems. Its unifying principle is the continuous conversion of sparse or image-based confocal measurements into low-latency spatial corrections. The principal design variables—scan geometry, dwell time, detector modality, estimator, actuator bandwidth, and photophysical budget—set the achievable compromise among localization precision, temporal resolution, axial range, and measurement duration.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Real-Time Confocal Tracking.