Papers
Topics
Authors
Recent
Search
2000 character limit reached

Kinetic Super-Resolution in Dynamic Imaging

Updated 12 July 2026
  • Kinetic super-resolution is a framework where resolution enhancement is integrated with dynamic models, enabling joint recovery of positions, velocities, and physical states from low-frequency, multitime data.
  • It employs techniques such as phase-space lifting, motion-aware reconstruction, and physics-informed constraints to overcome limitations of traditional static interpolation and tracking.
  • Numerical experiments and PDE-constrained methods demonstrate improved accuracy, spectral fidelity, and computational efficiency in recovering turbulent flows and dynamic imaging applications.

Kinetic super-resolution designates a family of inverse problems in which resolution enhancement is coupled to motion, temporal evolution, or governing dynamics rather than treated as purely spatial interpolation. In the literature summarized here, the term covers recovery of moving point sources from low-frequency multitime measurements, motion-aware reconstruction from subpixel translations or motion blur, and coarse-to-fine reconstruction of turbulent or PDE states using physics-informed or generative models (Alberti et al., 2018, Litterio et al., 21 May 2025, Geiping et al., 2016, Chehaitly, 15 Feb 2025, Kelshaw et al., 2022, Page, 2024, Sarkar et al., 2023, Shamooni et al., 4 Jul 2025, Schiødt et al., 19 Aug 2025). A related but distinct usage appears in interferometric phase estimation, where “super-resolution” refers to narrower phase fringes rather than spatiotemporal state recovery; that distinction is explicit in the phase-measurement literature (Schäfermeier et al., 2017).

1. Scope and defining characteristics

Across these works, kinetic super-resolution is characterized by the direct use of kinematic or dynamical structure inside the reconstruction map. Motion may appear as a latent variable to be estimated jointly with position, as a known measurement operator induced by sensor trajectory, as temporal consistency constraints between high-resolution frames, or as a PDE-constrained evolution law on the reconstructed state. This suggests that the adjective “kinetic” refers less to a single algorithm than to a class of formulations in which time dependence is part of the inverse problem itself.

A concise typology appears below.

Setting Core formulation Representative papers
Moving sparse sources Recover (x,v,w)(x,v,w) jointly in phase space from multitime low-frequency data (Alberti et al., 2018)
Motion-aware image/video SR Use subpixel translations, motion blur, or inter-frame optical flow as reconstruction constraints (Litterio et al., 21 May 2025, Geiping et al., 2016, Chehaitly, 15 Feb 2025)
Dynamics-aware PDE/turbulence SR Infer fine fields from coarse observations using PDE residuals, solver unrolling, or generative conditional models (Kelshaw et al., 2022, Page, 2024, Sarkar et al., 2023, Shamooni et al., 4 Jul 2025, Schiødt et al., 19 Aug 2025)

The same literature also emphasizes what kinetic super-resolution is not. It is not merely geometric upsampling, because the target is often a physically admissible high-resolution state rather than a visually sharper field. It is also not identical to conventional super-resolution defined from artificially downsampled ground truth. In CFD-oriented work, this distinction is central: coarse solver outputs are not equivalent to downsampled fine solutions, because downsampling preserves more of the underlying physics than an authentic coarse-grid simulation (Sarkar et al., 2023).

2. Phase-space lifting and dynamic spike super-resolution

A mathematically explicit formulation of kinetic super-resolution appears in the moving-spike setting. The unknown signal is a time-varying measure

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],

with discrete observation times

tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,

and low-frequency measurements

yk=Fμtk.y_k = F\mu_{t_k}.

In the low-frequency Fourier case,

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.

The key idea is to lift the problem from physical space to phase space (x,v)(x,v), replacing framewise reconstruction plus tracking by a single sparse inverse problem on

Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},

with sparse measure

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.

The reconstruction program becomes

minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}

This replaces the standard two-step pipeline—static localization at each time, then tracking—by simultaneous recovery of positions, velocities, and weights (Alberti et al., 2018).

The theory is organized around dynamical dual certificates of the form

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}

Exact recovery follows when μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],0 and μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],1 away from the support. A notable obstruction is the existence of ghost particles, defined as alternative trajectories that pass through the same observed locations at all sampled times. Theorem 5 states that if each frame is sufficiently separated so that a static dual certificate exists and the configuration has no ghost particles, then the true dynamic measure is the unique solution of the dynamical TV problem. Proposition 7 shows that, for particles drawn independently from absolutely continuous distributions, ghost particles are absent with probability 1. The noisy discrete setting is also treated: in 1D, under μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],2 and a discrete analogue of the no-ghost condition, Theorem 10 yields

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],3

with

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],4

The numerical experiments use 1D low-frequency Fourier measurements with μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],5 and μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],6 (five frames). They report that dynamical reconstruction succeeds more often than static reconstruction, especially when particles are close and static recovery is difficult. The same framework is then applied to ultrafast ultrasound localization microscopy, where microbubbles act as moving point reflectors; the proposed workflow selects time intervals with approximately constant total μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],7-norm and applies the dynamical reconstruction to each interval, yielding super-resolved vessel imaging together with blood-flow velocity estimation (Alberti et al., 2018).

3. Structured motion as a measurement resource

A second branch of kinetic super-resolution treats motion itself as an informative measurement operator. In structured-motion imaging, the low-resolution data are acquired under known sensor motion, either as a grid of subpixel translations or as continuous motion during exposure. For a moving sensor trajectory μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],8, the recorded low-resolution image satisfies

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],9

The moving-sensor image can be represented as a convolution with an occupancy map tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,0 induced by the motion. With super-resolution factor tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,1, interlacing all tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,2 subpixel shifts yields

tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,3

so the problem reduces to deconvolution with a box filter. For moving-sensor data the model becomes

tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,4

The central claim is that motion blur is not necessarily a nuisance: with high-precision motion information, sparse image priors, and convex optimization, it can aid reconstruction. Numerical experiments show that pseudo-random motion can reconstruct a high-resolution target from a single low-resolution image, and the reported real-data system works up to super-resolution factor tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,5 (Litterio et al., 21 May 2025).

In video super-resolution, motion can be built directly into the high-resolution reconstruction rather than used only for pairwise frame alignment. A variational model jointly reconstructs a batch of high-resolution frames tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,6 from low-resolution frames tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,7:

tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,8

The temporal and spatial regularizers are coupled through the motion-compensated discrepancy

tk=kT,k=K,,K,t_k = kT,\qquad k=-K,\dots,K,9

and the scale yk=Fμtk.y_k = F\mu_{t_k}.0 is chosen automatically. The essential consequence is linear, rather than quadratic, growth in the number of motion-estimation problems, because only neighboring-frame flows are required. On the reported yk=Fμtk.y_k = F\mu_{t_k}.1 benchmark, the proposed method attains 29.19 dB average PSNR, compared with 29.13 dB for VDSR, and it reaches 23.97 dB on the synthetic tube example (Geiping et al., 2016).

Classical multi-frame formulations remain important in this lineage. A standard observation model writes each low-resolution frame as

yk=Fμtk.y_k = F\mu_{t_k}.2

or equivalently yk=Fμtk.y_k = F\mu_{t_k}.3, where yk=Fμtk.y_k = F\mu_{t_k}.4 is the motion matrix, yk=Fμtk.y_k = F\mu_{t_k}.5 the blur operator, and yk=Fμtk.y_k = F\mu_{t_k}.6 the decimation matrix. One thesis-level treatment combines Horn–Schunck optical flow with TV-regularized reconstruction,

yk=Fμtk.y_k = F\mu_{t_k}.7

and solves the resulting problem by majorization-minimization with conjugate-gradient inner steps. The method yields visually improved high-resolution outputs from multiple yk=Fμtk.y_k = F\mu_{t_k}.8 low-resolution frames, but the implementation is explicitly described as not real-time, with reported processing around 1.5 hours in MATLAB experiments (Chehaitly, 15 Feb 2025).

4. Physics-informed reconstruction from sparse or coarse observations

In dynamical systems and turbulence, kinetic super-resolution is often defined by the requirement that reconstructed fine scales remain consistent with the governing PDE. One formulation learns a map

yk=Fμtk.y_k = F\mu_{t_k}.9

from sparse observations on a coarse grid (Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.0 to a high-resolution state on (Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.1, while minimizing a combined observation and residual loss:

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.2

The observation term enforces agreement with the known low-resolution samples, whereas the residual term penalizes violation of

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.3

For Kolmogorov flow, the governing system is the 2D incompressible Navier–Stokes equation

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.4

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.5

with

(Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.6

The implementation uses a VDSR architecture with bi-cubic upsampling, residual learning, periodic padding, a differentiable pseudospectral discretization, forward Euler time integration with (Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.7, and residual evaluation in the Fourier domain. The reported setup uses (Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.8, (Fμt)=[0,1]de2πixdμt(x),fc.(F\mu_t)_\ell = \int_{[0,1]^d} e^{-2\pi i x\cdot \ell}\,d\mu_t(x), \qquad |\ell|\le f_c.9, 2048 training windows, 256 validation windows, (x,v)(x,v)0, Adam, learning rate (x,v)(x,v)1, and (x,v)(x,v)2. Average relative (x,v)(x,v)3 errors are 0.0872 for the physics-informed CNN, 0.2091 for bi-linear interpolation, and 0.1717 for bi-cubic interpolation (Kelshaw et al., 2022).

A distinct but related strategy puts dynamics directly into the training loss through a differentiable solver. For forced two-dimensional turbulence on the torus, the governing equation is

(x,v)(x,v)4

with velocity recovered from

(x,v)(x,v)5

Instead of supervising the network by a high-resolution target at (x,v)(x,v)6, the method trains an initial condition whose forward evolution remains consistent with observations. The coarse-only loss is

(x,v)(x,v)7

This requires a fully differentiable solver in the loop; the implementation uses JAX-CFD with a pseudospectral solver, high-resolution simulations on (x,v)(x,v)8 for (x,v)(x,v)9 and Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},0 for Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},1, forcing mode Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},2, and Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},3 at Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},4. The model has reconstruction errors similar to standard supervised super-resolution despite using no high-resolution reference data. At Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},5 and coarse factor Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},6, the coarse-only error at Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},7 is only about Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},8 the high-resolution-trained model’s error, and after forward evolution the errors become comparable to, and sometimes smaller than, those from high-resolution-trained models. The same work reports that the learned model outperforms variational data assimilation for initial state estimation, with assimilation errors exceeding 40% at the harsher Q={(x,v)R2d: x+kTv[0,1]d, k[K,K]},\mathcal{Q} = \{(x,v)\in\mathbb{R}^{2d}:\ x+kTv\in[0,1]^d,\ \forall k\in[-K,K]\},9 setting (Page, 2024).

These formulations treat super-resolution as recovery of missing physics rather than filling in between pixels. In both cases, temporal evolution supplies the admissibility criterion for the fine-scale reconstruction.

5. Coarse-grid to fine-grid PDE prediction

Another major formulation redefines super-resolution specifically for PDE-based computation. The mapping is written as

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.0

where w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.1 is a coarse-grid CFD solution and w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.2 is a fine-grid CFD solution. The defining claim is that coarse inputs should come from actual coarse-grid simulations rather than from artificial downsampling of fine data. A physics-infused UNet is used to learn the nonlinear relationship between coarse mesh data

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.3

and fine mesh data

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.4

The total loss is

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.5

where the physics term is an MSE over convective and diffusive or conductive derivative terms computed by 2nd-order finite differences (Sarkar et al., 2023).

The architecture consists of four contracting convolution blocks, a bottleneck, four expansion blocks with skip connections, a w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.6 output projection, and a bilinear upsampler. The formulation is demonstrated on 2D Burgers’ equation, methane combustion, and an industrial heat exchanger. For Burgers’ equation,

w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.7

the reported grid mapping is w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.8. For methane combustion the mapping is w=i=1Nwiδ(xi,vi).w=\sum_{i=1}^N w_i\,\delta_{(x_i,v_i)}.9, and for the heat exchanger it is minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}0. The Burgers case uses

minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}1

while the methane and heat-exchanger cases use analogous weighted convective and diffusive or conductive penalties (Sarkar et al., 2023).

The reported numerical results are explicit. For Burgers’ minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}2, bilinear interpolation has RMSE 0.4927, bicubic 0.5133, UNet 0.0283, and PIUNet 0.0207; for minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}3, PIUNet reaches RMSE 0.0225 and minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}4. The paper also reports mean minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}5 error of PIUNet on the fine mesh as 0.0097, compared with coarse-mesh minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}6 error 0.0446, corresponding to roughly 78% reduction. In methane combustion, the adiabatic flame temperature RMSE decreases from 78.317 for bilinear and 77.374 for bicubic to 20.954 for PIUNet, with about 73% improvement in minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}7 RMSE versus baseline interpolation; methane mass fraction RMSE improves by about 20.47%. In the heat-exchanger case, PIUNet reduces RMSE approximately by 36.8% for minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}8, 30.7% for minνM(Q)νTVsubject toGν=y.(6)\min_{\nu\in\mathcal{M}(\mathcal{Q})} \|\nu\|_{\mathrm{TV}} \quad\text{subject to}\quad G\nu = y. \tag{6}9, 35.04% for q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}0, and 20.06% for q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}1. The corresponding coarse-solver-plus-network workflow yields about 24× speedup for Burgers, about 93× for methane combustion, and about 235× for the heat exchanger (Sarkar et al., 2023).

This line of work frames kinetic super-resolution as a surrogate fine-mesh solver. The output is not a photographic refinement but an approximate fine-grid PDE state intended to preserve convective transport, diffusion or conduction, and boundary-condition structure.

6. Turbulence, subgrid physics, and generative reconstruction

In turbulence, kinetic super-resolution is often assessed by whether it reconstructs spectra, vorticity statistics, dissipation, and localized high-frequency structure. One particle-aware formulation uses a conditional GAN for two-way coupled particle-laden turbulent flows. The generator reconstructs the high-resolution velocity field as

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}2

where

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}3

is subgrid kinetic energy and

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}4

is effective particle mass density. The discriminator is conditioned on the low-resolution input and stationary-wavelet-transform detail coefficients,

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}5

with

q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}6

The generator uses a deep RRDB network with 16 RRDB blocks, each with 3 RDBs; the discriminator is a U-Net with spectral normalization. Training and testing use DNS from forced homogeneous isotropic turbulence and decaying turbulence, with q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}7, q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}8, q(x,v)=k=KKfcck,e2πi(x+kTv).(9)q(x,v)=\sum_{k=-K}^K\sum_{|\ell|\le f_c} c_{k,\ell}\,e^{2\pi i \ell\cdot(x+kTv)}. \tag{9}9 slices, 16,000 training samples per case, 800 test samples per case, total training set 320,000 samples, and total test set 16,000 samples. The ablation that masks particle conditioning at inference shows that large-scale structures remain somewhat reasonable but high-wavenumber energy decays too quickly and vorticity structures become less accurate. Including particle input improves μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],00 from μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],01 to μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],02, stress-tensor error from μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],03 to μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],04, and velocity error from μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],05 to μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],06. The same model reproduces both positive subgrid dissipation and backscatter, in contrast to the standard Smagorinsky closure

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],07

which is always positive (Shamooni et al., 4 Jul 2025).

A generative alternative uses stochastic interpolants for two-dimensional turbulence. The target field μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],08 is sampled on a μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],09 grid, while the coarse input is obtained by low-pass filtering with μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],10 and downsampling to μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],11. The conditional generative task is to learn

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],12

with interpolant

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],13

and SDE

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],14

For the turbulence application,

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],15

Inference uses a Heun SDE integrator with 100 pseudo-timesteps. The main methodological extension is a patch-wise strategy: the μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],16 field is split into 16 patches of size μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],17, reconstructed in two stages by a free-generator and a cond-generator arranged in a checkerboard-like procedure. This reduces patch-boundary artifacts, especially in the vorticity field. The reported diagnostics include the kinetic energy spectrum, vorticity,

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],18

local dissipation

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],19

and PDFs of spatially averaged kinetic energy, vorticity skewness, and spatially averaged dissipation. For dissipation, the base field has KL divergence 0.5480, while SI reduces it to 0.0106 or 0.0040 depending on the variant; Wasserstein-1 drops from 2.5809 to 0.3013 or 0.1094. The paper states that stochastic interpolants outperform flow-matching and diffusion baselines across a range of metrics, and that the patch-wise approach often matches or exceeds full-field reconstruction (Schiødt et al., 19 Aug 2025).

Taken together, these results place subgrid physics at the center of kinetic super-resolution. The relevant target is not only pointwise accuracy but the recovery of unresolved energy, dissipation, intermittency, and particle-modulated fine-scale structure.

7. Conceptual distinctions, misconceptions, and recurring limitations

Several recurring distinctions structure the field. First, dynamic or kinetic super-resolution is repeatedly contrasted with static reconstruction followed by tracking. In the moving-spike literature, the standard pipeline fails when particles are too close in any single frame, ignores temporal information during reconstruction, and requires a separate tracking stage; phase-space lifting is proposed specifically to avoid those drawbacks (Alberti et al., 2018).

Second, motion is not uniformly treated as degradation. Structured-motion imaging argues that motion blur can be helpful for super-resolution and that pseudo-random motion can encode usable spatial information in a single blurred exposure (Litterio et al., 21 May 2025). Video super-resolution similarly uses motion to couple unknown high-resolution frames directly, rather than merely transporting low-resolution evidence frame by frame (Geiping et al., 2016). At the same time, classical motion-based methods remain limited by motion-estimation accuracy, brightness constancy assumptions, occlusions, and computational cost (Chehaitly, 15 Feb 2025, Geiping et al., 2016).

Third, several PDE-oriented works reject the assumption that super-resolution should be trained on downsampled fine solutions. The coarse CFD literature explicitly states that downsampling fine data is not the correct analogue for PDE super-resolution, because true coarse-grid simulations embody different physics (Sarkar et al., 2023). A related misconception is that turbulent super-resolution is equivalent to image sharpening. Physics-informed and dynamics-in-the-loss methods instead define success through PDE consistency, spectral fidelity, or future evolution under a solver (Kelshaw et al., 2022, Page, 2024).

Finally, the term “super-resolution” itself is field-dependent. In deterministic phase measurements, super-resolution means interference fringes narrower than the usual half-wavelength periodicity, while super-sensitivity means phase uncertainty below the shot-noise limit,

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],20

That work uses a coherent state, a squeezed vacuum state, homodyne detection, and a dichotomic windowing strategy

μt=i=1Nwiδxi+vit,t[δ,δ],\mu_t = \sum_{i=1}^N w_i\,\delta_{x_i+v_i t}, \qquad t\in[-\delta,\delta],21

to achieve both effects simultaneously, reporting 430 photons, a 22-fold improvement in phase resolution, and a 1.7-fold improvement in sensitivity. The paper’s distinction between fringe width and estimation precision clarifies that “super-resolution” is not a universal synonym for better inference; its technical meaning depends on the measurement problem being posed (Schäfermeier et al., 2017).

In the kinetic literature proper, a plausible implication is that future progress will continue to depend on embedding the correct dynamics into the inverse problem. The most successful formulations here do not simply increase pixel density. They reconstruct trajectories, exploit structured motion, or infer fine-scale states constrained by PDEs, spectra, and multiscale statistics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Kinetic Super-Resolution.