Kinetic Super-Resolution in Dynamic Imaging
- Kinetic super-resolution is a framework where resolution enhancement is integrated with dynamic models, enabling joint recovery of positions, velocities, and physical states from low-frequency, multitime data.
- It employs techniques such as phase-space lifting, motion-aware reconstruction, and physics-informed constraints to overcome limitations of traditional static interpolation and tracking.
- Numerical experiments and PDE-constrained methods demonstrate improved accuracy, spectral fidelity, and computational efficiency in recovering turbulent flows and dynamic imaging applications.
Kinetic super-resolution designates a family of inverse problems in which resolution enhancement is coupled to motion, temporal evolution, or governing dynamics rather than treated as purely spatial interpolation. In the literature summarized here, the term covers recovery of moving point sources from low-frequency multitime measurements, motion-aware reconstruction from subpixel translations or motion blur, and coarse-to-fine reconstruction of turbulent or PDE states using physics-informed or generative models (Alberti et al., 2018, Litterio et al., 21 May 2025, Geiping et al., 2016, Chehaitly, 15 Feb 2025, Kelshaw et al., 2022, Page, 2024, Sarkar et al., 2023, Shamooni et al., 4 Jul 2025, Schiødt et al., 19 Aug 2025). A related but distinct usage appears in interferometric phase estimation, where “super-resolution” refers to narrower phase fringes rather than spatiotemporal state recovery; that distinction is explicit in the phase-measurement literature (Schäfermeier et al., 2017).
1. Scope and defining characteristics
Across these works, kinetic super-resolution is characterized by the direct use of kinematic or dynamical structure inside the reconstruction map. Motion may appear as a latent variable to be estimated jointly with position, as a known measurement operator induced by sensor trajectory, as temporal consistency constraints between high-resolution frames, or as a PDE-constrained evolution law on the reconstructed state. This suggests that the adjective “kinetic” refers less to a single algorithm than to a class of formulations in which time dependence is part of the inverse problem itself.
A concise typology appears below.
| Setting | Core formulation | Representative papers |
|---|---|---|
| Moving sparse sources | Recover jointly in phase space from multitime low-frequency data | (Alberti et al., 2018) |
| Motion-aware image/video SR | Use subpixel translations, motion blur, or inter-frame optical flow as reconstruction constraints | (Litterio et al., 21 May 2025, Geiping et al., 2016, Chehaitly, 15 Feb 2025) |
| Dynamics-aware PDE/turbulence SR | Infer fine fields from coarse observations using PDE residuals, solver unrolling, or generative conditional models | (Kelshaw et al., 2022, Page, 2024, Sarkar et al., 2023, Shamooni et al., 4 Jul 2025, Schiødt et al., 19 Aug 2025) |
The same literature also emphasizes what kinetic super-resolution is not. It is not merely geometric upsampling, because the target is often a physically admissible high-resolution state rather than a visually sharper field. It is also not identical to conventional super-resolution defined from artificially downsampled ground truth. In CFD-oriented work, this distinction is central: coarse solver outputs are not equivalent to downsampled fine solutions, because downsampling preserves more of the underlying physics than an authentic coarse-grid simulation (Sarkar et al., 2023).
2. Phase-space lifting and dynamic spike super-resolution
A mathematically explicit formulation of kinetic super-resolution appears in the moving-spike setting. The unknown signal is a time-varying measure
with discrete observation times
and low-frequency measurements
In the low-frequency Fourier case,
The key idea is to lift the problem from physical space to phase space , replacing framewise reconstruction plus tracking by a single sparse inverse problem on
with sparse measure
The reconstruction program becomes
This replaces the standard two-step pipeline—static localization at each time, then tracking—by simultaneous recovery of positions, velocities, and weights (Alberti et al., 2018).
The theory is organized around dynamical dual certificates of the form
Exact recovery follows when 0 and 1 away from the support. A notable obstruction is the existence of ghost particles, defined as alternative trajectories that pass through the same observed locations at all sampled times. Theorem 5 states that if each frame is sufficiently separated so that a static dual certificate exists and the configuration has no ghost particles, then the true dynamic measure is the unique solution of the dynamical TV problem. Proposition 7 shows that, for particles drawn independently from absolutely continuous distributions, ghost particles are absent with probability 1. The noisy discrete setting is also treated: in 1D, under 2 and a discrete analogue of the no-ghost condition, Theorem 10 yields
3
with
4
The numerical experiments use 1D low-frequency Fourier measurements with 5 and 6 (five frames). They report that dynamical reconstruction succeeds more often than static reconstruction, especially when particles are close and static recovery is difficult. The same framework is then applied to ultrafast ultrasound localization microscopy, where microbubbles act as moving point reflectors; the proposed workflow selects time intervals with approximately constant total 7-norm and applies the dynamical reconstruction to each interval, yielding super-resolved vessel imaging together with blood-flow velocity estimation (Alberti et al., 2018).
3. Structured motion as a measurement resource
A second branch of kinetic super-resolution treats motion itself as an informative measurement operator. In structured-motion imaging, the low-resolution data are acquired under known sensor motion, either as a grid of subpixel translations or as continuous motion during exposure. For a moving sensor trajectory 8, the recorded low-resolution image satisfies
9
The moving-sensor image can be represented as a convolution with an occupancy map 0 induced by the motion. With super-resolution factor 1, interlacing all 2 subpixel shifts yields
3
so the problem reduces to deconvolution with a box filter. For moving-sensor data the model becomes
4
The central claim is that motion blur is not necessarily a nuisance: with high-precision motion information, sparse image priors, and convex optimization, it can aid reconstruction. Numerical experiments show that pseudo-random motion can reconstruct a high-resolution target from a single low-resolution image, and the reported real-data system works up to super-resolution factor 5 (Litterio et al., 21 May 2025).
In video super-resolution, motion can be built directly into the high-resolution reconstruction rather than used only for pairwise frame alignment. A variational model jointly reconstructs a batch of high-resolution frames 6 from low-resolution frames 7:
8
The temporal and spatial regularizers are coupled through the motion-compensated discrepancy
9
and the scale 0 is chosen automatically. The essential consequence is linear, rather than quadratic, growth in the number of motion-estimation problems, because only neighboring-frame flows are required. On the reported 1 benchmark, the proposed method attains 29.19 dB average PSNR, compared with 29.13 dB for VDSR, and it reaches 23.97 dB on the synthetic tube example (Geiping et al., 2016).
Classical multi-frame formulations remain important in this lineage. A standard observation model writes each low-resolution frame as
2
or equivalently 3, where 4 is the motion matrix, 5 the blur operator, and 6 the decimation matrix. One thesis-level treatment combines Horn–Schunck optical flow with TV-regularized reconstruction,
7
and solves the resulting problem by majorization-minimization with conjugate-gradient inner steps. The method yields visually improved high-resolution outputs from multiple 8 low-resolution frames, but the implementation is explicitly described as not real-time, with reported processing around 1.5 hours in MATLAB experiments (Chehaitly, 15 Feb 2025).
4. Physics-informed reconstruction from sparse or coarse observations
In dynamical systems and turbulence, kinetic super-resolution is often defined by the requirement that reconstructed fine scales remain consistent with the governing PDE. One formulation learns a map
9
from sparse observations on a coarse grid 0 to a high-resolution state on 1, while minimizing a combined observation and residual loss:
2
The observation term enforces agreement with the known low-resolution samples, whereas the residual term penalizes violation of
3
For Kolmogorov flow, the governing system is the 2D incompressible Navier–Stokes equation
4
5
with
6
The implementation uses a VDSR architecture with bi-cubic upsampling, residual learning, periodic padding, a differentiable pseudospectral discretization, forward Euler time integration with 7, and residual evaluation in the Fourier domain. The reported setup uses 8, 9, 2048 training windows, 256 validation windows, 0, Adam, learning rate 1, and 2. Average relative 3 errors are 0.0872 for the physics-informed CNN, 0.2091 for bi-linear interpolation, and 0.1717 for bi-cubic interpolation (Kelshaw et al., 2022).
A distinct but related strategy puts dynamics directly into the training loss through a differentiable solver. For forced two-dimensional turbulence on the torus, the governing equation is
4
with velocity recovered from
5
Instead of supervising the network by a high-resolution target at 6, the method trains an initial condition whose forward evolution remains consistent with observations. The coarse-only loss is
7
This requires a fully differentiable solver in the loop; the implementation uses JAX-CFD with a pseudospectral solver, high-resolution simulations on 8 for 9 and 0 for 1, forcing mode 2, and 3 at 4. The model has reconstruction errors similar to standard supervised super-resolution despite using no high-resolution reference data. At 5 and coarse factor 6, the coarse-only error at 7 is only about 8 the high-resolution-trained model’s error, and after forward evolution the errors become comparable to, and sometimes smaller than, those from high-resolution-trained models. The same work reports that the learned model outperforms variational data assimilation for initial state estimation, with assimilation errors exceeding 40% at the harsher 9 setting (Page, 2024).
These formulations treat super-resolution as recovery of missing physics rather than filling in between pixels. In both cases, temporal evolution supplies the admissibility criterion for the fine-scale reconstruction.
5. Coarse-grid to fine-grid PDE prediction
Another major formulation redefines super-resolution specifically for PDE-based computation. The mapping is written as
0
where 1 is a coarse-grid CFD solution and 2 is a fine-grid CFD solution. The defining claim is that coarse inputs should come from actual coarse-grid simulations rather than from artificial downsampling of fine data. A physics-infused UNet is used to learn the nonlinear relationship between coarse mesh data
3
and fine mesh data
4
The total loss is
5
where the physics term is an MSE over convective and diffusive or conductive derivative terms computed by 2nd-order finite differences (Sarkar et al., 2023).
The architecture consists of four contracting convolution blocks, a bottleneck, four expansion blocks with skip connections, a 6 output projection, and a bilinear upsampler. The formulation is demonstrated on 2D Burgers’ equation, methane combustion, and an industrial heat exchanger. For Burgers’ equation,
7
the reported grid mapping is 8. For methane combustion the mapping is 9, and for the heat exchanger it is 0. The Burgers case uses
1
while the methane and heat-exchanger cases use analogous weighted convective and diffusive or conductive penalties (Sarkar et al., 2023).
The reported numerical results are explicit. For Burgers’ 2, bilinear interpolation has RMSE 0.4927, bicubic 0.5133, UNet 0.0283, and PIUNet 0.0207; for 3, PIUNet reaches RMSE 0.0225 and 4. The paper also reports mean 5 error of PIUNet on the fine mesh as 0.0097, compared with coarse-mesh 6 error 0.0446, corresponding to roughly 78% reduction. In methane combustion, the adiabatic flame temperature RMSE decreases from 78.317 for bilinear and 77.374 for bicubic to 20.954 for PIUNet, with about 73% improvement in 7 RMSE versus baseline interpolation; methane mass fraction RMSE improves by about 20.47%. In the heat-exchanger case, PIUNet reduces RMSE approximately by 36.8% for 8, 30.7% for 9, 35.04% for 0, and 20.06% for 1. The corresponding coarse-solver-plus-network workflow yields about 24× speedup for Burgers, about 93× for methane combustion, and about 235× for the heat exchanger (Sarkar et al., 2023).
This line of work frames kinetic super-resolution as a surrogate fine-mesh solver. The output is not a photographic refinement but an approximate fine-grid PDE state intended to preserve convective transport, diffusion or conduction, and boundary-condition structure.
6. Turbulence, subgrid physics, and generative reconstruction
In turbulence, kinetic super-resolution is often assessed by whether it reconstructs spectra, vorticity statistics, dissipation, and localized high-frequency structure. One particle-aware formulation uses a conditional GAN for two-way coupled particle-laden turbulent flows. The generator reconstructs the high-resolution velocity field as
2
where
3
is subgrid kinetic energy and
4
is effective particle mass density. The discriminator is conditioned on the low-resolution input and stationary-wavelet-transform detail coefficients,
5
with
6
The generator uses a deep RRDB network with 16 RRDB blocks, each with 3 RDBs; the discriminator is a U-Net with spectral normalization. Training and testing use DNS from forced homogeneous isotropic turbulence and decaying turbulence, with 7, 8, 9 slices, 16,000 training samples per case, 800 test samples per case, total training set 320,000 samples, and total test set 16,000 samples. The ablation that masks particle conditioning at inference shows that large-scale structures remain somewhat reasonable but high-wavenumber energy decays too quickly and vorticity structures become less accurate. Including particle input improves 00 from 01 to 02, stress-tensor error from 03 to 04, and velocity error from 05 to 06. The same model reproduces both positive subgrid dissipation and backscatter, in contrast to the standard Smagorinsky closure
07
which is always positive (Shamooni et al., 4 Jul 2025).
A generative alternative uses stochastic interpolants for two-dimensional turbulence. The target field 08 is sampled on a 09 grid, while the coarse input is obtained by low-pass filtering with 10 and downsampling to 11. The conditional generative task is to learn
12
with interpolant
13
and SDE
14
For the turbulence application,
15
Inference uses a Heun SDE integrator with 100 pseudo-timesteps. The main methodological extension is a patch-wise strategy: the 16 field is split into 16 patches of size 17, reconstructed in two stages by a free-generator and a cond-generator arranged in a checkerboard-like procedure. This reduces patch-boundary artifacts, especially in the vorticity field. The reported diagnostics include the kinetic energy spectrum, vorticity,
18
local dissipation
19
and PDFs of spatially averaged kinetic energy, vorticity skewness, and spatially averaged dissipation. For dissipation, the base field has KL divergence 0.5480, while SI reduces it to 0.0106 or 0.0040 depending on the variant; Wasserstein-1 drops from 2.5809 to 0.3013 or 0.1094. The paper states that stochastic interpolants outperform flow-matching and diffusion baselines across a range of metrics, and that the patch-wise approach often matches or exceeds full-field reconstruction (Schiødt et al., 19 Aug 2025).
Taken together, these results place subgrid physics at the center of kinetic super-resolution. The relevant target is not only pointwise accuracy but the recovery of unresolved energy, dissipation, intermittency, and particle-modulated fine-scale structure.
7. Conceptual distinctions, misconceptions, and recurring limitations
Several recurring distinctions structure the field. First, dynamic or kinetic super-resolution is repeatedly contrasted with static reconstruction followed by tracking. In the moving-spike literature, the standard pipeline fails when particles are too close in any single frame, ignores temporal information during reconstruction, and requires a separate tracking stage; phase-space lifting is proposed specifically to avoid those drawbacks (Alberti et al., 2018).
Second, motion is not uniformly treated as degradation. Structured-motion imaging argues that motion blur can be helpful for super-resolution and that pseudo-random motion can encode usable spatial information in a single blurred exposure (Litterio et al., 21 May 2025). Video super-resolution similarly uses motion to couple unknown high-resolution frames directly, rather than merely transporting low-resolution evidence frame by frame (Geiping et al., 2016). At the same time, classical motion-based methods remain limited by motion-estimation accuracy, brightness constancy assumptions, occlusions, and computational cost (Chehaitly, 15 Feb 2025, Geiping et al., 2016).
Third, several PDE-oriented works reject the assumption that super-resolution should be trained on downsampled fine solutions. The coarse CFD literature explicitly states that downsampling fine data is not the correct analogue for PDE super-resolution, because true coarse-grid simulations embody different physics (Sarkar et al., 2023). A related misconception is that turbulent super-resolution is equivalent to image sharpening. Physics-informed and dynamics-in-the-loss methods instead define success through PDE consistency, spectral fidelity, or future evolution under a solver (Kelshaw et al., 2022, Page, 2024).
Finally, the term “super-resolution” itself is field-dependent. In deterministic phase measurements, super-resolution means interference fringes narrower than the usual half-wavelength periodicity, while super-sensitivity means phase uncertainty below the shot-noise limit,
20
That work uses a coherent state, a squeezed vacuum state, homodyne detection, and a dichotomic windowing strategy
21
to achieve both effects simultaneously, reporting 430 photons, a 22-fold improvement in phase resolution, and a 1.7-fold improvement in sensitivity. The paper’s distinction between fringe width and estimation precision clarifies that “super-resolution” is not a universal synonym for better inference; its technical meaning depends on the measurement problem being posed (Schäfermeier et al., 2017).
In the kinetic literature proper, a plausible implication is that future progress will continue to depend on embedding the correct dynamics into the inverse problem. The most successful formulations here do not simply increase pixel density. They reconstruct trajectories, exploit structured motion, or infer fine-scale states constrained by PDEs, spectra, and multiscale statistics.