Discrete Scale-Space Filters
- Discrete scale-space filters are operators on sampled data that generate progressively coarser representations while rigorously preserving axiomatic properties such as semigroup structure and causality.
- They extend classical Gaussian smoothing by incorporating non-Gaussian linear constructions, inverse scale-space methods, and learned receptive fields to enhance detail control and robustness.
- These filters include advanced strategies like time-causal recursive architectures and optimized digital filter designs that enable efficient, multiscale analysis in modern imaging and signal processing.
Discrete scale-space filters are operators on sampled signals, images, or other discrete data that generate families of progressively coarser representations indexed by a discrete or discretized scale parameter while preserving scale-space structure as far as the discrete domain allows. In the classical linear setting, the reference model is Gaussian smoothing and its discrete diffusion analogues: continuous scale-space theory singles out the Gaussian under linearity, shift invariance, normalization, and non-enhancement of local extrema, while discrete theory distinguishes between mere grid-sampled Gaussian approximations and genuinely discrete kernels that preserve semigroup and causality properties on the lattice (Kubota, 2011, Lindeberg, 2023). Subsequent work has broadened the topic in several directions: non-Gaussian linear constructions with hidden state, time-causal recursive temporal filters, nonlinear inverse-scale-space iterations, stochastic scale-spaces, morphological and steerable multiscale architectures, and learned receptive fields that can be modeled by discrete scale-space primitives (Lindeberg, 2017, Bednarski et al., 2022, Lindeberg et al., 16 Sep 2025).
1. Classical axioms and the canonical discrete Gaussian
Classical scale-space theory asks for a family of operators that simplify structure with increasing scale without introducing spurious detail. In the continuous domain, this yields Gaussian smoothing; in the discrete domain, the same question is more delicate because sampled data live on a lattice, scales are often discrete, and naive sampling of the continuous Gaussian does not automatically preserve the defining axioms. The central genuinely discrete smoothing kernel is the discrete analogue of the Gaussian,
with the modified Bessel function of integer order. This kernel is normalized, has variance , satisfies the exact semigroup property
and arises from the semi-discrete diffusion equation, which makes it the axiomatic discrete counterpart of Gaussian scale-space (Lindeberg, 2023).
A common misconception is that any sampled Gaussian on a grid is already a valid discrete scale-space filter. The discrete literature separates three different objects: sampled Gaussian kernels, integrated Gaussian kernels, and the discrete Gaussian analogue. Sampled Gaussian and sampled Gaussian derivative kernels approximate the continuous theory well when the scale parameter is sufficiently large; the cited experiments state that this occurs when the scale parameter is greater than a value of about $1$, in units of the grid spacing. At very fine scales, however, sampled and integrated Gaussian derivatives perform very poorly, whereas the discrete analogue with central differences remains substantially better behaved and exactly consistent with polynomial tests in exact arithmetic (Lindeberg, 2023).
In the older scalar iterative viewpoint, repeated application of a single convolution kernel has frequency response
Under the scale-space constraints recalled in the literature, this forces Gaussian-type diffusion behavior; in the discrete case this reduces to the discrete diffusion or binomial family. Kubota’s formulation makes this point explicit by contrasting the scalar case with later multichannel constructions (Kubota, 2011).
2. Beyond Gaussianity in linear discrete systems
The strongest linear departure from classical Gaussianity in the supplied literature is the matrix-of-filters construction of Kubota. Instead of iterating a single scalar kernel, the method evolves a primary signal together with an auxiliary hidden signal through a fixed matrix of local filters,
with the observed output obtained by projecting the multichannel state back to the primary component. The effective primary transfer function is then no longer a pure power law. Rather, it becomes
a frequencywise convex combination of powers of two eigenmodes (Kubota, 2011).
This change in functional form is the mechanism by which non-Gaussian linear scale space becomes possible. The paper derives necessary structural conditions and a rigorous sufficient parameter region in terms of guaranteeing real, nonnegative, normalized, unimodal, and monotonically reducing responses for all 0. The parameter 1 is decisive: when 2, the construction collapses to the Gaussian or discrete-diffusion case; when 3, genuinely non-Gaussian behavior appears. The explicit example
4
produces a compact-support non-Gaussian scale space whose frequency response develops a sharper cutoff and narrower transition than standard diffusion while preserving positivity, monotone reduction, and unimodality (Kubota, 2011).
An important conceptual consequence follows. The Gaussian uniqueness result applies to single convolution-kernel constructions, not to broader linear iterative systems with hidden state. The visible output of the 5 system is the projection of a multichannel linear dynamical system, not the repeated application of a single scalar kernel. This establishes a new class of discrete linear scale-space filters without contradicting the classical scalar uniqueness theorem (Kubota, 2011).
3. Discretization, approximation, and filter design
Discrete scale-space filtering is not only a question of axioms; it is also a question of approximation quality and implementability. A substantial part of the recent literature compares explicit discrete Gaussian approximations, smoothing-plus-difference schemes, and constrained filter design. One branch compares sampled Gaussian kernels, integrated Gaussian kernels, and the discrete Gaussian analogue as smoothing stages, then studies derivative computation either by explicit derivative kernels or by central differences after smoothing. In this taxonomy, the “hybrid discretizations” use either the normalized sampled Gaussian or the integrated Gaussian followed by small-support central differences. They are computationally attractive when multiple derivative orders at the same scale are required, because the expensive smoothing stage is shared, but they have larger offsets from the continuous theory than explicit sampled or integrated derivative kernels, and the genuinely discrete method based on 6 usually matches the intended amount of smoothing better (Lindeberg, 2024).
The same comparison sharpens the interpretation of “fine scale.” For first through fourth derivatives, hybrid methods inherit lower bounds from the derivative stencils themselves, such as
7
so their effective spread cannot shrink arbitrarily as the nominal scale decreases. This is one reason why very small scales behave differently across discretizations and why the discrete analogue is preferred when the interpretation of scale at fine resolution matters (Lindeberg, 2024).
A more engineering-oriented line is represented by pyDOF, which provides a constrained optimization framework for symmetric physical-space discrete filters. Its transfer-function constraints include exact DC preservation, positivity, boundedness in 8, and monotonic attenuation, through keywords such as fixOne, zeroToOne, and monotoneNeg. This supports the design of compact low-pass filters with scale-space-like spectral surrogates, including Gaussian targets, but the paper is explicit that it does not prove classical scale-space properties such as non-enhancement of local extrema, semigroup structure across scales, or positivity of the spatial kernel coefficients (Nikolaou et al., 25 Jun 2026).
A related but distinct alternative is the direct digital design of multiscale blur-plus-differentiator filters with vanishing moments. “Digital filters with vanishing moments for shape analysis” derives separable FIR and IIR filters via derivative constraints at DC, recommends a two-stage low-pass/high-pass architecture, and gives explicit repeated-pole IIR blur coefficients as a function of scale. This framework is presented as a digital-domain alternative to Gaussian-derived filters: it is highly relevant to discrete multiscale filtering, but it prioritizes vanishing moments, separability, and computational efficiency rather than classical scale-space axioms (Kennedy, 2019).
4. Time-causal and recursive temporal scale-space filters
Temporal scale-space filtering introduces an additional structural requirement: causality. For continuous video or audio streams, future samples are unavailable, so temporal filters must satisfy 9 for 0 and ideally be time-recursive. Lindeberg’s time-causal construction uses cascades of first-order integrators with primitive kernel
1
and composed transfer function
2
Each added stage defines a new temporal scale level, so the resulting temporal scale-space family is discrete by construction (Lindeberg, 2017).
Two scale distributions are contrasted. With equal time constants, the temporal variance is 3, so the scales are uniform in 4. This works reasonably for sparse peaks and onset ramps, but it does not yield true scale invariance, and sine-wave scale estimates are biased. The alternative uses a logarithmic distribution
5
which leads, in the infinite-cascade limit, to the scale-invariant time-causal limit kernel with Fourier transform
6
This kernel is closed under temporal rescalings 7, 8, and therefore transfers a large part of Gaussian scale-selection theory to a genuinely causal recursive setting (Lindeberg, 2017).
Scale selection is performed through local extrema over temporal scale of scale-normalized derivatives. For the limit kernel, the selected scale transforms covariantly,
9
and for 0 the corresponding scale-normalized magnitudes are invariant. The result is technically notable: true temporal scale invariance is achievable although the temporal scale levels have to be discrete (Lindeberg, 2017).
5. Nonlinear, variational, stochastic, and geometric generalizations
Discrete scale-space filtering is not confined to linear convolutional smoothing. Inverse scale-space methods reinterpret “scale” as the iteration parameter of a nonlinear variational process. For convex absolutely one-homogeneous regularizers, the inverse scale-space flow
1
is discretized by Bregman iteration,
2
Early iterates are coarse; later ones progressively add detail. The 2022 lifting-based extension shows how this idea can be generalized to non-convex data terms by running the Bregman iteration on a convex lifted functional with sublabel-accurate discretization. In the anisotropic-TV case, the paper proves a subgradient transformation theorem ensuring that the discrete lifted iteration reduces to the standard Bregman iteration when the lifted subgradient has the factorized form 3, which preserves the interpretation as a genuine nonlinear discrete scale-space filter (Bednarski et al., 2022).
A different generalization treats diffusion probabilistic models as stochastic scale-spaces. Here the forward process is a discrete-time Markov chain on image-valued random variables,
4
so the object evolving with scale is not a single image but a probability distribution over images. The paper establishes analogues of initial state, semigroup structure, monotone simplification via entropy, and asymptotic steady states, then connects blur-based variants to the discrete heat semigroup 5 and to osmosis filtering. This is a scale-space in probability space rather than a classical image-space convolution scale-space, but it makes the relation between modern diffusion models and classical multiscale filtering mathematically explicit (Peter, 2023).
At the geometric-processing boundary of the subject, log-aesthetic surface filtering supplies an iterative local nonlinear mesh filter that steers a triangular mesh toward a target Gaussian-curvature distribution. The implemented filter restricts a vertex update to
6
estimates Gaussian curvature by angle deficit, fits a local plane to neighboring curvature values, and uses bisection to solve for 7. This is not a canonical discrete scale-space construction in the diffusion or semigroup sense, but it is an iterative discrete smoothing filter with scale-space-like progressive fairing behavior (Miura et al., 2013).
6. Learned, equivariant, and anisotropic multiscale filter banks
Recent work has also moved discrete scale-space filters into learned architectures. One line begins from the observation that scale-equivariant processing on discrete images requires a scale-space lifting before semigroup cross-correlation. The framework based on semigroup cross-correlation over discrete scales and translations allows Gaussian, morphological dilation and erosion, and morphological opening and closing as admissible liftings. In experiments, all scale-space liftings outperform a CNN baseline for generalization to unseen scales, and on the geometric segmentation task over scales 8, quadratic dilation reached 9 IoU, compared with $1$0 for Gaussian lifting (Sangalli et al., 2021).
A second line embeds Gaussian derivative scale-space directly into convolutional kernels. The Rotation-Scale Equivariant Steerable Filter represents each learned kernel as a linear combination of sampled Gaussian derivative atoms,
$1$1
with trainable scale parameters constrained to disjoint intervals across scale groups. This imports classical Gaussian derivative receptive fields and steerability into a discrete CNN layer, producing an approximately rotation-scale equivariant filter family rather than an arbitrary learned pixel kernel (Yang et al., 2023).
The descriptive power of discrete scale-space models for learned filters is pushed further by the “master key filters” analysis of depthwise-separable networks. There, clustered ConvNeXt receptive fields are modeled by small-support difference operators applied to the discrete analogue of the Gaussian kernel, including one-sided derivatives, centered derivatives, Gaussian smoothing, and Laplacian-based sharpening. The empirical replacement experiment is unusually strong: Original ConvNeXt v2 Tiny reaches $1$2 Top-1 accuracy on ImageNet, Frozen 8 master key filters also reach $1$3, Frozen 8 filters from Method B reach $1$4, and allowing learning of the scale parameters gives $1$5. This shows that learned depthwise filters in that architecture can be well approximated by discrete scale-space filters (Lindeberg et al., 16 Sep 2025).
Not all multiscale filter banks in the literature are scale spaces in the classical Gaussian sense. The anisotropic QMF construction for wavelet and generalized shearlet systems uses anisotropic dilation matrices, Smith factorization, perfect reconstruction, and vanishing moments to build critically sampled orthogonal filter banks. These are rigorously multiscale and discrete, but they are better regarded as anisotropic multiresolution filter banks than as causal Gaussian-type scale-space filters (Cotronei et al., 2018). This distinction matters: discrete scale-space filtering and discrete multiscale filtering overlap substantially, but they are not identical categories.