Papers
Topics
Authors
Recent
Search
2000 character limit reached

OPE-Upscale Module: Orthogonal Image SR

Updated 30 June 2026
  • The OPE-Upscale module is a parameter-free upsampling mechanism that uses orthogonal position encoding to reconstruct high-resolution images for arbitrary scales.
  • It replaces conventional INR-based modules with a deterministic process employing fixed trigonometric evaluations and matrix-vector multiplications to ensure efficient, mathematically interpretable inference.
  • Empirical evaluations demonstrate competitive fidelity with state-of-the-art methods while significantly reducing computational and memory overhead, enabling faster super-resolution rendering.

The OPE-Upscale module is a parameter-free upsampling mechanism designed for arbitrary-scale image super-resolution (SR). It replaces conventional implicit neural representation (INR)-based upsampling modules by leveraging orthogonal position encoding (OPE). The OPE-Upscale module reconstructs high-resolution images in a mathematically interpretable and efficient manner, achieving competitive fidelity with state-of-the-art approaches while significantly reducing computational and memory requirements (Song et al., 2023).

1. Orthogonal Position Encoding: Mathematical Formulation

The foundation of the OPE-Upscale module is the orthogonal position encoding (OPE), an extension of standard position encoding. OPE defines an explicit, orthonormal basis for mapping 2D coordinates within [−1,1]2[-1,1]^2 to a high-dimensional embedding. For an input coordinate (x,y)(x, y) and maximum frequency n∈Nn \in \mathbb{N}:

  • The 1D position encoding map is

γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .

This yields a vector in R2n+1\mathbb{R}^{2n+1}.

  • The 2D encoding P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2} is constructed by the outer product and flattening:

X=γ(x),Y=γ(y),P(x,y)=flatten(XTY)X = \gamma(x),\quad Y = \gamma(y),\quad P(x, y) = \text{flatten}(X^T Y)

  • The local reconstruction of a continuous image channel f(x,y)f(x, y) is expressed as

f(x,y)≈ZP(x,y)Tf(x, y) \approx ZP(x, y)^T

where ZZ is a learned coefficient (“projection”) vector.

This encoding forms an orthonormal basis under the (x,y)(x, y)0 inner product, with explicit expressions for all basis functions (combining (x,y)(x, y)1 and (x,y)(x, y)2 terms in both directions).

2. OPE-Upscale Module: Architecture and Rendering Procedure

The OPE-Upscale module is structured around a clear separation of learning and deterministic inference:

  • Encoder (x,y)(x, y)3: A conventional convolutional network (e.g., EDSR-baseline or RDN) processes the low-resolution image (x,y)(x, y)4 to generate the feature map (x,y)(x, y)5. Each spatial location (x,y)(x, y)6 in (x,y)(x, y)7 contains the concatenated coefficient vectors (x,y)(x, y)8, (x,y)(x, y)9, n∈Nn \in \mathbb{N}0 for the RGB channels.
  • Rendering at Arbitrary Grid: For each output pixel coordinate n∈Nn \in \mathbb{N}1:

    1. Locate the nearest feature-map cell center n∈Nn \in \mathbb{N}2.
    2. Compute local relative coordinates:

    n∈Nn \in \mathbb{N}3

  1. Form n∈Nn \in \mathbb{N}4 using OPE described above.
  2. Reconstruct the SR pixel value for each channel via

    n∈Nn \in \mathbb{N}5

  3. To ensure seamless stitching, a weighted patch-ensemble of the four nearest neighbors (using bilinear interpolation weights) is used:

    n∈Nn \in \mathbb{N}6

This process efficiently handles arbitrary-scale and continuous coordinates.

3. Parameter-Free and Analytical Properties

The OPE-Upscale module is characterized by its complete absence of trainable parameters in the upsampling stage:

  • All learned parameters are contained within the encoder n∈Nn \in \mathbb{N}7.
  • The upsampling pipeline consists solely of fixed trigonometric evaluations, matrix-vector products, and linear combinations without any neural network layers (such as MLPs or convolutions) in the module itself.
  • Given a feature map n∈Nn \in \mathbb{N}8, every SR pixel is deterministically computed, establishing the OPE-Upscale module as analytically interpretable and fully parameter-free at inference.

4. Orthonormality and Mathematical Justification

OPE’s encoding functions form a mathematically orthonormal basis within the finite domain n∈Nn \in \mathbb{N}9 under the γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .0 inner product:

γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .1

The family of basis functions consists of combinations of

  • γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .2
  • γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .3
  • γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .4
  • γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .5

for γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .6, with normalization factors to ensure orthonormality. The orthogonality can be demonstrated by direct calculation of the relevant inner products—after incorporating the γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .7 scaling in γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .8, it follows that

γ(x)=[1,2cos⁡(πx),2sin⁡(πx),…,2cos⁡(nπx),2sin⁡(nπx)] .\gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .9

where R2n+1\mathbb{R}^{2n+1}0 denote the corresponding basis functions. The encoding R2n+1\mathbb{R}^{2n+1}1 thus provides an orthonormal expansion suitable for analytical super-resolution reconstruction.

5. Algorithmic Workflow

The rendering algorithm for a high-resolution image is as follows:

P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}8

This procedure leverages only cos/sin evaluations and matrix-vector multiplications per pixel for highly efficient rendering.

6. Empirical Evaluation and Resource Analysis

Extensive experimentation confirms the following properties:

  • Fidelity: On DIV2K-val (arbitrary scales R2n+1\mathbb{R}^{2n+1}2–R2n+1\mathbb{R}^{2n+1}3), the OPE-SR method narrows the PSNR gap to LIIF/LTE to less than 0.1 dB in most cases. On standard benchmarks (Set5, Set14, B100, Urban100), the drop is less than 0.15 dB. For extreme super-resolution factors (R2n+1\mathbb{R}^{2n+1}4–R2n+1\mathbb{R}^{2n+1}5), OPE matches or outperforms competitors.
  • Efficiency: The computational requirements per SR pixel are approximately R2n+1\mathbb{R}^{2n+1}6 multiply–accumulates and a handful of trigonometric function calls, compared to R2n+1\mathbb{R}^{2n+1}7 in LIIF. Overall FLOPs for a full image are R2n+1\mathbb{R}^{2n+1}8 million versus R2n+1\mathbb{R}^{2n+1}9 billion. OPE-Upscale achieves P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}0–P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}1 times faster system inference (e.g., EDSR+LIIF P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}2 s/image vs. EDSR+OPE P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}3 s/image on DIV2K-val); rendering alone is P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}4–P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}5\% faster on large images.
  • Memory Use: The module uses zero additional activations or gradients during training. LIIF/LTE incur P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}6–P(x,y)∈R(2n+1)2P(x, y) \in \mathbb{R}^{(2n+1)^2}7 MB of memory overhead, while OPE-Upscale incurs virtually none.

These results establish that the OPE-Upscale module enables mathematically interpretable, parameter-free, and resource-efficient arbitrary-scale image super-resolution while maintaining competitive output quality (Song et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to OPE-Upscale Module.