---
title: Anisotropic Fourier Feature Positional Encoding
url: https://www.emergentmind.com/topics/anisotropic-fourier-feature-positional-encoding-afpe
type: topic
---

# Anisotropic Fourier Feature Positional Encoding

Anisotropic Fourier Feature Positional Encoding (AFPE) is a positional encoding scheme introduced for medical imaging that generalizes isotropic Fourier Feature Positional Encoding (IFPE) by replacing a single isotropic frequency scale with dimension-specific scales. In the formulation proposed in “Anisotropic Fourier Features for Positional Encoding in Medical Imaging” [2509.02488], AFPE is intended to align positional similarity with anisotropic image geometry, target anatomy, and spatiotemporal structure rather than treating all coordinate directions as metrically equivalent. The paper argues that positional encoding is consequential rather than incidental in medical transformers, and reports that the optimal positional encoding depends on the shape of the structure of interest and the anisotropy of the data; AFPE is reported to significantly outperform state-of-the-art positional encodings in all tested anisotropic settings [2509.02488].

## 1. Problem setting and motivation

AFPE is motivated by the observation that medical imaging is frequently high-dimensional and anisotropic. The paper emphasizes three recurring forms of anisotropy: 3D volumetric anisotropy, where MRI and CT often have different in-plane and through-plane resolution; directional anatomical structure, where organs, vessels, and disease patterns are often elongated or direction-dependent rather than isotropic; and spatiotemporal anisotropy in videos, where time is not metrically interchangeable with spatial axes [2509.02488].

Within this setting, Transformers require explicit positional information because attention is permutation-invariant. The paper argues that standard sinusoidal positional encodings (SPEs), inherited from language modeling and extended to images by encoding each axis separately and concatenating them, can be suboptimal in higher-dimensional geometric domains. It further argues that IFPE improves distance modeling relative to SPE, but remains mismatched when spatial axes have unequal physical spacing, when temporal and spatial axes belong to different metric regimes, or when the target structure itself has directional bias [2509.02488].

This motivation is consistent with a broader Fourier-feature perspective in which the frequency distribution used to construct positional features determines the geometry seen by the model. Tancik et al. analyze Fourier feature mappings as a way to transform the effective NTK into a stationary kernel with tunable bandwidth, making the choice of frequency distribution the central design parameter [2006.10739]. AFPE inherits that premise but removes the isotropy assumption at the level of positional scale selection.

## 2. Formal definition

The paper defines a positional encoding as a function
\[
f:p\rightarrow \mathbb{R}^D
\]
mapping an \(m\)-dimensional position \(p\) to an embedding vector
\[
d=(d_1,d_2,\dots,d_D) \in R^D,
\]
where \(D\) matches the token embedding dimension [2509.02488].

For reference, the paper gives the sinusoidal baseline as
\[
SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \\ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}
\]
with temperature \(t\) and \(freq_j = \frac{2d_j}{D}\). It gives isotropic Fourier Feature Positional Encoding as
\[
IFPE_s(p) = [\sin(2\pi B_s p) || \cos(2\pi B_s p)]^T ,
\]
where
\[
B_s \in \mathbb{R}^{m \times D/2}
\]
is sampled from
\[
B_s \sim \mathcal{N}(0,s).
\]
The learnable Fourier baseline is
\[
LFPE(p)=[\sin(2 \pi Wp)  || \cos(2\pi Wp)]^T,
\]
where
\[
W \in \mathbb{R}^{m \times D/2}
\]
is trainable and initialized by
\[
W \sim \mathcal{N}(0, 1)
\]
[2509.02488].

AFPE is obtained by replacing the single isotropic scale with per-dimension scales. The exact definition given is
\[
B_{s_i} \sim \mathcal{N}\left(0, s_i\right) \quad \text{for all } i \in \{1, \dots, m\},
\]
where
\[
s = (s_1, s_2, \dots, s_m)^\top \in \mathbb{R}^m
\]
and each \(s_i > 0\) is the scale for the \(i\)-th spatial dimension. This Gaussian is then used in the same Fourier mapping as IFPE:
\[
AFPE(p) = [\sin(2\pi B_s p) || \cos(2\pi B_s p)]^T .
\]
The paper does not introduce a full covariance matrix or non-diagonal anisotropic Gaussian; its anisotropy model is strictly axis-wise scaling [2509.02488].

The construction is presented generically for \(m\)-dimensional coordinates. For 2D images, \(m=2\); for 3D volumes, \(m=3\); and for spatiotemporal video, \(m=3\) with one temporal and two spatial coordinates. In all cases, the procedure is the same: sample or construct \(B_s\) according to per-axis scales, compute \(2\pi B_s p\), concatenate sine and cosine features to obtain a \(D\)-dimensional positional embedding, and add the embedding elementwise to patch tokens in a ViT pipeline [2509.02488].

## 3. Geometric interpretation and anisotropic inductive bias

The geometric point of AFPE is that not all coordinate directions should be treated equally. In IFPE, positional similarity is isotropic because one scale hyperparameter affects all axes equally. In AFPE, each axis has its own scale, so similarity can decay differently along different directions. The paper presents this as the ability to model stronger continuity along one axis, weaker continuity along another, and distinct spatial versus temporal similarity regimes in video [2509.02488].

The paper repeatedly connects this directional control to anatomical and pathological shape priors. It argues that elongated organs or patterns may benefit from anisotropic continuity, whereas round structures may favor isotropic settings. In the ChestX analysis, \(s_{\text{row} < s_{\text{column}}\) works well for Atelectasis, Pneumothorax, Cardiomegaly, Tortuous Aorta, Subcutaneous Emphysema, and Pneumomediastinum; \(s_{\text{row} > s_{\text{column}}\) works well for Emphysema, Pleural Thickening, Infiltration, Effusion, and Nodule; and \(s_{\text{row} = s_{\text{column}}\) is reported as optimal for Mass, which the paper links to its round shape [2509.02488].

For EchoNet, the final selected values are
\[
s_{\text{time}=0.785,\qquad s_{\text{space}=0.111,
\]
with the two spatial dimensions constrained to share the same scale. The paper interprets this as reflecting greater independence across time than across space [2509.02488].

A plausible interpretation is that AFPE replaces isotropic geometry with an axis-weighted geometry. The paper states that its definition is equivalent to a diagonal-covariance Gaussian over the projection vectors, but does not write the full multivariate distribution explicitly. This suggests that AFPE can be read as a stationary Fourier-feature encoding whose correlation structure differs by axis rather than remaining uniform in every direction [2509.02488].

This axis-sensitive view is related to learnable Fourier-feature formulations in which projection vectors define orientation and wavelength. “Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding” describes a trainable matrix \(W_r\) whose rows define both orientation and wavelength of Fourier features, but there anisotropy is emergent because \(W_r\) is unconstrained and learned rather than specified by explicit per-axis scales [2106.02795].

## 4. Relation to adjacent positional encodings and common misconceptions

AFPE is best understood against three neighboring families: classical sinusoidal encodings, isotropic Fourier-feature encodings, and learnable Fourier-feature encodings. SPE encodes each axis separately and is criticized in the AFPE paper for struggling to preserve Euclidean distances in higher-dimensional spaces. IFPE improves multi-dimensional distance modeling, but assumes a single shared scale. LFPE makes the projection matrix trainable, but does not explicitly encode domain anisotropy through fixed axis-specific scales [2509.02488].

The general theoretical background comes from Fourier-feature work rather than from medical imaging alone. Tancik et al. show that Fourier feature mappings enable MLPs to learn high-frequency functions in low-dimensional domains, and characterize their effect through a stationary kernel with tunable bandwidth [2006.10739]. AFPE differs by shifting the design question from a single global bandwidth to a vector of per-dimension scales.

Several nearby arXiv papers are AFPE-relevant without actually proposing AFPE. “3QFP: Efficient neural implicit surface reconstruction using Tri-Quadtrees and Fourier feature Positional encoding” uses a standard Gaussian Fourier feature positional encoding with coefficients sampled from an isotropic Gaussian distribution; its Fourier encoding is explicitly isotropic, and any directional structure enters through tri-planar feature storage rather than through the Fourier map itself [2401.07164]. “Rethinking Positional Encoding for Neural Vehicle Routing” proposes a hierarchical anisometric positional encoding built from distance-indexed sinusoidal in-route encoding and depot-anchored angular cross-route encoding; it is Fourier-like and anisometric, but it is not the medical-imaging AFPE formulation and does not use AFPE as an official term [2605.11910].

A common misconception is therefore to treat any directional or geometry-aware sinusoidal encoding as AFPE. In the strict sense established by [2509.02488], AFPE denotes an anisotropic generalization of IFPE in which dimension-specific Gaussian scales are used to sample Fourier projections for medical-image coordinates.

## 5. Empirical results

The paper evaluates AFPE on five tasks: multi-label classification on ChestX, organ classification on OrganMNIST3D, ejection-fraction regression on EchoNet Dynamic, and Feret-diameter regression on AdrenalMNIST3D and VesselMNIST3D. The compared positional encodings are None, Learnable, SPE, IFPE, LFPE, and AFPE [2509.02488].

The main quantitative results support a conditional claim rather than a universal one. In isotropic settings, AFPE is competitive but not always uniquely best; in anisotropic settings, AFPE degrades least and is typically best. The most pronounced improvement appears in spatiotemporal echocardiography, where temporal and spatial dimensions are explicitly decoupled [2509.02488].

| Task / setting | Metric | Key result |
|---|---:|---|
| ChestX | AUPRC | AFPE \(0.185 \pm 0.004\); SPE \(0.186 \pm 0.003\); IFPE \(0.184 \pm 0.003\) |
| OrganMNIST3D, anisotropy 1 | AUROC | AFPE \(0.994 \pm 0.001\); SPE \(0.995 \pm 0.001\); IFPE \(0.994 \pm 0.001\) |
| OrganMNIST3D, anisotropy 4 | AUROC | AFPE \(0.984 \pm 0.002\); SPE \(0.980 \pm 0.003\); IFPE \(0.979 \pm 0.003\) |
| OrganMNIST3D, anisotropy 6 | AUROC | AFPE \(0.983 \pm 0.002\); SPE \(0.975 \pm 0.003\); IFPE \(0.975 \pm 0.003\) |
| OrganMNIST3D, anisotropy 8 | AUROC | AFPE \(0.972 \pm 0.003\); SPE \(0.969 \pm 0.003\); IFPE \(0.958 \pm 0.004\) |
| EchoNet, anisotropy 7 | \(R^2\) | AFPE \(0.621 \pm 0.024\); IFPE \(0.547 \pm 0.027\); SPE \(0.527 \pm 0.028\) |

The EchoNet result is the strongest single result in the paper. At anisotropy 7, AFPE reaches \(0.621 \pm 0.024\) in \(R^2\), compared with \(0.547 \pm 0.027\) for IFPE, \(0.527 \pm 0.028\) for SPE, \(0.522 \pm 0.029\) for LFPE, \(0.334 \pm 0.029\) for Learnable, and \(0.283 \pm 0.034\) for no positional encoding [2509.02488].

The Feret-diameter ablation is important because it separates generic shape encoding from anisotropy handling. On AdrenalMNIST3D and VesselMNIST3D, Fourier-based encodings massively outperform SPE for shape-descriptor learning, while IFPE and AFPE are broadly similar. For example, on AdrenalMNIST3D at anisotropy 3, SPE yields \(0.071 \pm 0.067\) for Feret minimum and \(0.750 \pm 0.083\) for Feret maximum, whereas IFPE yields \(0.725 \pm 0.032\) and \(0.921 \pm 0.010\), and AFPE yields \(0.683 \pm 0.031\) and \(0.922 \pm 0.013\). The paper interprets this as evidence that AFPE’s largest gains arise from anisotropy modeling rather than from a generic superiority in shape representation [2509.02488].

## 6. Implementation, limitations, and future directions

AFPE is introduced as a replacement for the positional encoding function rather than as an architectural change to the Transformer. It is added to ViT patch tokens in the usual way, so parameter count is effectively unchanged except for negligible positional-encoding specification, and training or inference complexity is unchanged relative to other fixed absolute positional encodings [2509.02488].

The method is mostly fixed and hyperparameterized rather than learned end-to-end. For ChestX and EchoNet, the paper reports 50 training runs with randomly sampled \(s_i\), each trained for 75 epochs, after which the best scales were selected on validation performance. For OrganMNIST3D, no hyperparameter optimization was done; instead,
\[
s_i = 0.5 \times \text{anisotropy}_i.
\]
The paper presents this as a direct practical heuristic when acquisition anisotropy is known [2509.02488].

The limitations are equally explicit. First, AFPE is hand-specified or tuned rather than learned automatically. Second, anisotropy is modeled only axis-wise: the formulation uses separate \(s_i\) per dimension but no full covariance matrix, rotation-aware anisotropy, or non-axis-aligned metric. Third, the benefits are strongest in anisotropic settings rather than universally dominant across all tasks. Fourth, the theoretical development is limited: the paper motivates AFPE geometrically and empirically but does not provide a full new kernel derivation or formal proof of anisotropic distance preservation [2509.02488].

The paper proposes future work on segmentation models such as SAM, object detection frameworks such as DETR, and end-to-end learning of scale parameters. It also states that the formulation itself is general even though the experiments are medical: beyond medical imaging, anisotropic coordinate structure also appears in videos, remote sensing, scientific simulation grids, and robotics spatiotemporal data. This suggests a broader relevance for AFPE, but the paper does not experimentally validate those domains [2509.02488].

Source: https://www.emergentmind.com/topics/anisotropic-fourier-feature-positional-encoding-afpe