OPE-Upscale Module: Orthogonal Image SR
- The OPE-Upscale module is a parameter-free upsampling mechanism that uses orthogonal position encoding to reconstruct high-resolution images for arbitrary scales.
- It replaces conventional INR-based modules with a deterministic process employing fixed trigonometric evaluations and matrix-vector multiplications to ensure efficient, mathematically interpretable inference.
- Empirical evaluations demonstrate competitive fidelity with state-of-the-art methods while significantly reducing computational and memory overhead, enabling faster super-resolution rendering.
The OPE-Upscale module is a parameter-free upsampling mechanism designed for arbitrary-scale image super-resolution (SR). It replaces conventional implicit neural representation (INR)-based upsampling modules by leveraging orthogonal position encoding (OPE). The OPE-Upscale module reconstructs high-resolution images in a mathematically interpretable and efficient manner, achieving competitive fidelity with state-of-the-art approaches while significantly reducing computational and memory requirements (Song et al., 2023).
1. Orthogonal Position Encoding: Mathematical Formulation
The foundation of the OPE-Upscale module is the orthogonal position encoding (OPE), an extension of standard position encoding. OPE defines an explicit, orthonormal basis for mapping 2D coordinates within to a high-dimensional embedding. For an input coordinate and maximum frequency :
- The 1D position encoding map is
This yields a vector in .
- The 2D encoding is constructed by the outer product and flattening:
- The local reconstruction of a continuous image channel is expressed as
where is a learned coefficient (“projection”) vector.
This encoding forms an orthonormal basis under the 0 inner product, with explicit expressions for all basis functions (combining 1 and 2 terms in both directions).
2. OPE-Upscale Module: Architecture and Rendering Procedure
The OPE-Upscale module is structured around a clear separation of learning and deterministic inference:
- Encoder 3: A conventional convolutional network (e.g., EDSR-baseline or RDN) processes the low-resolution image 4 to generate the feature map 5. Each spatial location 6 in 7 contains the concatenated coefficient vectors 8, 9, 0 for the RGB channels.
- Rendering at Arbitrary Grid: For each output pixel coordinate 1:
- Locate the nearest feature-map cell center 2.
- Compute local relative coordinates:
3
- Form 4 using OPE described above.
- Reconstruct the SR pixel value for each channel via
5
- To ensure seamless stitching, a weighted patch-ensemble of the four nearest neighbors (using bilinear interpolation weights) is used:
6
This process efficiently handles arbitrary-scale and continuous coordinates.
3. Parameter-Free and Analytical Properties
The OPE-Upscale module is characterized by its complete absence of trainable parameters in the upsampling stage:
- All learned parameters are contained within the encoder 7.
- The upsampling pipeline consists solely of fixed trigonometric evaluations, matrix-vector products, and linear combinations without any neural network layers (such as MLPs or convolutions) in the module itself.
- Given a feature map 8, every SR pixel is deterministically computed, establishing the OPE-Upscale module as analytically interpretable and fully parameter-free at inference.
4. Orthonormality and Mathematical Justification
OPE’s encoding functions form a mathematically orthonormal basis within the finite domain 9 under the 0 inner product:
1
The family of basis functions consists of combinations of
- 2
- 3
- 4
- 5
for 6, with normalization factors to ensure orthonormality. The orthogonality can be demonstrated by direct calculation of the relevant inner products—after incorporating the 7 scaling in 8, it follows that
9
where 0 denote the corresponding basis functions. The encoding 1 thus provides an orthonormal expansion suitable for analytical super-resolution reconstruction.
5. Algorithmic Workflow
The rendering algorithm for a high-resolution image is as follows:
8
This procedure leverages only cos/sin evaluations and matrix-vector multiplications per pixel for highly efficient rendering.
6. Empirical Evaluation and Resource Analysis
Extensive experimentation confirms the following properties:
- Fidelity: On DIV2K-val (arbitrary scales 2–3), the OPE-SR method narrows the PSNR gap to LIIF/LTE to less than 0.1 dB in most cases. On standard benchmarks (Set5, Set14, B100, Urban100), the drop is less than 0.15 dB. For extreme super-resolution factors (4–5), OPE matches or outperforms competitors.
- Efficiency: The computational requirements per SR pixel are approximately 6 multiply–accumulates and a handful of trigonometric function calls, compared to 7 in LIIF. Overall FLOPs for a full image are 8 million versus 9 billion. OPE-Upscale achieves 0–1 times faster system inference (e.g., EDSR+LIIF 2 s/image vs. EDSR+OPE 3 s/image on DIV2K-val); rendering alone is 4–5\% faster on large images.
- Memory Use: The module uses zero additional activations or gradients during training. LIIF/LTE incur 6–7 MB of memory overhead, while OPE-Upscale incurs virtually none.
These results establish that the OPE-Upscale module enables mathematically interpretable, parameter-free, and resource-efficient arbitrary-scale image super-resolution while maintaining competitive output quality (Song et al., 2023).