---
title: 'OPE-Upscale Module: Orthogonal Image SR'
url: https://www.emergentmind.com/topics/ope-upscale-module
type: topic
---

# OPE-Upscale Module: Orthogonal Image SR

The OPE-Upscale module is a parameter-free upsampling mechanism designed for arbitrary-scale image super-resolution (SR). It replaces conventional implicit neural representation (INR)-based upsampling modules by leveraging orthogonal position encoding (OPE). The OPE-Upscale module reconstructs high-resolution images in a mathematically interpretable and efficient manner, achieving competitive fidelity with state-of-the-art approaches while significantly reducing computational and memory requirements [2303.01091].

## 1. Orthogonal Position Encoding: Mathematical Formulation

The foundation of the OPE-Upscale module is the orthogonal position encoding (OPE), an extension of standard position encoding. OPE defines an explicit, orthonormal basis for mapping 2D coordinates within $[-1,1]^2$ to a high-dimensional embedding. For an input coordinate $(x, y)$ and maximum frequency $n \in \mathbb{N}$:

- The 1D position encoding map is
  $$
  \gamma(x) = [1, \sqrt{2}\cos(\pi x), \sqrt{2}\sin(\pi x), \ldots, \sqrt{2}\cos(n\pi x), \sqrt{2}\sin(n\pi x)]\ .
  $$
  This yields a vector in $\mathbb{R}^{2n+1}$.

- The 2D encoding $P(x, y) \in \mathbb{R}^{(2n+1)^2}$ is constructed by the outer product and flattening:
  $$
  X = \gamma(x),\quad Y = \gamma(y),\quad P(x, y) = \text{flatten}(X^T Y)
  $$

- The local reconstruction of a continuous image channel $f(x, y)$ is expressed as
  $$
  f(x, y) \approx ZP(x, y)^T
  $$
  where $Z$ is a learned coefficient (“projection”) vector.

This encoding forms an orthonormal basis under the $L^2([−1,1]^2)$ inner product, with explicit expressions for all basis functions (combining $\cos$ and $\sin$ terms in both directions).

## 2. OPE-Upscale Module: Architecture and Rendering Procedure

The OPE-Upscale module is structured around a clear separation of learning and deterministic inference:

- **Encoder $E_\theta$:** A conventional convolutional network (e.g., EDSR-baseline or RDN) processes the low-resolution image $I_{lr} \in \mathbb{R}^{h \times w \times 3}$ to generate the feature map $C \in \mathbb{R}^{h \times w \times 3(2n+1)^2}$. Each spatial location $(i, j)$ in $C$ contains the concatenated coefficient vectors ${\bf Z}_R$, ${\bf Z}_G$, ${\bf Z}_B$ for the RGB channels.

- **Rendering at Arbitrary Grid:** For each output pixel coordinate $(x_q, y_q) \in [-1, 1]^2$:
  1. Locate the nearest feature-map cell center $(x_c, y_c)$.
  2. Compute local relative coordinates:
     $$
     x' = (x_q - x_c) \cdot h,\quad y' = (y_q - y_c) \cdot w
     $$
  3. Form $P(x', y')$ using OPE described above.
  4. Reconstruct the SR pixel value for each channel via
     $$
     I_{SR}^R(x_q, y_q) = {\bf Z}_R P(x', y')^T\ , \text{ etc.}
     $$
  5. To ensure seamless stitching, a weighted patch-ensemble of the four nearest neighbors (using bilinear interpolation weights) is used:
     $$
     I_{SR}(x_q, y_q) = \sum_{t \in \{00, 01, 10, 11\}} \frac{s_t}{S} \cdot \mathcal{R}(z^*_t, (x_q'^t, y_q'^t))
     $$

This process efficiently handles arbitrary-scale and continuous coordinates.

## 3. Parameter-Free and Analytical Properties

The OPE-Upscale module is characterized by its complete absence of trainable parameters in the upsampling stage:

- All learned parameters are contained within the encoder $E_\theta$.
- The upsampling pipeline consists solely of fixed trigonometric evaluations, matrix-vector products, and linear combinations without any neural network layers (such as MLPs or convolutions) in the module itself.
- Given a feature map $C$, every SR pixel is deterministically computed, establishing the OPE-Upscale module as analytically interpretable and fully parameter-free at inference.

## 4. Orthonormality and Mathematical Justification

OPE’s encoding functions form a mathematically orthonormal basis within the finite domain $[-1, 1]^2$ under the $L^2$ inner product:
$$
\langle g, h \rangle = \frac{1}{4} \int_{-1}^{1} \int_{-1}^{1} g(x, y)h(x, y)\,dx\,dy\ .
$$

The family of basis functions consists of combinations of
- $\cos(k\pi x)\cos(\ell\pi y)$
- $\cos(k\pi x)\sin(\ell\pi y)$
- $\sin(k\pi x)\cos(\ell\pi y)$
- $\sin(k\pi x)\sin(\ell\pi y)$

for $k, \ell = 0,\ldots,n$, with normalization factors to ensure orthonormality. The orthogonality can be demonstrated by direct calculation of the relevant inner products—after incorporating the $\sqrt{2}$ scaling in $\gamma(\cdot)$, it follows that
$$
\langle e_{i_1, j_1}, e_{i_2, j_2} \rangle = \delta_{i_1, i_2} \delta_{j_1, j_2}
$$
where $e_{i,j}$ denote the corresponding basis functions. The encoding $P(x,y)$ thus provides an orthonormal expansion suitable for analytical super-resolution reconstruction.

## 5. Algorithmic Workflow

The rendering algorithm for a high-resolution image is as follows:

```python
# Inputs: feature_map C (h, w, 3*M), target resolution H, W, max frequency n
# Precompute position encodings
def gamma(t):
    G = [1]
    for k in range(1, n+1):
        G.append(sqrt(2) * cos(k * pi * t))
        G.append(sqrt(2) * sin(k * pi * t))
    return np.array(G)  # length 2n+1

def pos_enc_2d(x, y):
    X = gamma(x)
    Y = gamma(y)
    return X[:, None] * Y[None, :]  # Outer, then flatten to length M

for u in range(H):
    x_q = (u + 0.5)/H * 2 - 1
    for v in range(W):
        y_q = (v + 0.5)/W * 2 - 1
        # Compute bilinear weights and locate four nearest LR centers
        # For each neighbor t: compute relative coords, retrieve Z, form P, and sum weighted dot products
        # Aggregate pixel value for RGB
        pass  # Detailed steps follow as in section 2.
# Output: high-res image of size H x W x 3
```

This procedure leverages only cos/sin evaluations and matrix-vector multiplications per pixel for highly efficient rendering.

## 6. Empirical Evaluation and Resource Analysis

Extensive experimentation confirms the following properties:

- **Fidelity:** On DIV2K-val (arbitrary scales $2$–$30$), the OPE-SR method narrows the PSNR gap to LIIF/LTE to less than 0.1 dB in most cases. On standard benchmarks (Set5, Set14, B100, Urban100), the drop is less than 0.15 dB. For extreme super-resolution factors ($\times 6$–$\times 30$), OPE matches or outperforms competitors.

- **Efficiency:** The computational requirements per SR pixel are approximately $6\,000$ multiply–accumulates and a handful of trigonometric function calls, compared to $429\,000$ in LIIF. Overall FLOPs for a full image are $85$ million versus $6.2$ billion. OPE-Upscale achieves $2$–$3$ times faster system inference (e.g., EDSR+LIIF $\sim1.7$ s/image vs. EDSR+OPE $\sim0.48$ s/image on DIV2K-val); rendering alone is $26$–$57$\% faster on large images.

- **Memory Use:** The module uses zero additional activations or gradients during training. LIIF/LTE incur $85$–$98$ MB of memory overhead, while OPE-Upscale incurs virtually none.

These results establish that the OPE-Upscale module enables mathematically interpretable, parameter-free, and resource-efficient arbitrary-scale image super-resolution while maintaining competitive output quality [2303.01091].

Source: https://www.emergentmind.com/topics/ope-upscale-module