Anisotropic Fourier Feature Positional Encoding
- AFPE is a positional encoding method that uses per-dimension Fourier scales to capture anisotropic characteristics in medical imaging.
- It extends isotropic Fourier features by introducing axis-specific scales, aligning encoding with directional anatomical structures and data anisotropy.
- Empirical results demonstrate that AFPE outperforms standard encodings, especially in spatiotemporal tasks like echocardiography and organ classification.
Anisotropic Fourier Feature Positional Encoding (AFPE) is a positional encoding scheme introduced for medical imaging that generalizes isotropic Fourier Feature Positional Encoding (IFPE) by replacing a single isotropic frequency scale with dimension-specific scales. In the formulation proposed in “Anisotropic Fourier Features for Positional Encoding in Medical Imaging” (Jabareen et al., 2 Sep 2025), AFPE is intended to align positional similarity with anisotropic image geometry, target anatomy, and spatiotemporal structure rather than treating all coordinate directions as metrically equivalent. The paper argues that positional encoding is consequential rather than incidental in medical transformers, and reports that the optimal positional encoding depends on the shape of the structure of interest and the anisotropy of the data; AFPE is reported to significantly outperform state-of-the-art positional encodings in all tested anisotropic settings (Jabareen et al., 2 Sep 2025).
1. Problem setting and motivation
AFPE is motivated by the observation that medical imaging is frequently high-dimensional and anisotropic. The paper emphasizes three recurring forms of anisotropy: 3D volumetric anisotropy, where MRI and CT often have different in-plane and through-plane resolution; directional anatomical structure, where organs, vessels, and disease patterns are often elongated or direction-dependent rather than isotropic; and spatiotemporal anisotropy in videos, where time is not metrically interchangeable with spatial axes (Jabareen et al., 2 Sep 2025).
Within this setting, Transformers require explicit positional information because attention is permutation-invariant. The paper argues that standard sinusoidal positional encodings (SPEs), inherited from language modeling and extended to images by encoding each axis separately and concatenating them, can be suboptimal in higher-dimensional geometric domains. It further argues that IFPE improves distance modeling relative to SPE, but remains mismatched when spatial axes have unequal physical spacing, when temporal and spatial axes belong to different metric regimes, or when the target structure itself has directional bias (Jabareen et al., 2 Sep 2025).
This motivation is consistent with a broader Fourier-feature perspective in which the frequency distribution used to construct positional features determines the geometry seen by the model. Tancik et al. analyze Fourier feature mappings as a way to transform the effective NTK into a stationary kernel with tunable bandwidth, making the choice of frequency distribution the central design parameter (Tancik et al., 2020). AFPE inherits that premise but removes the isotropy assumption at the level of positional scale selection.
2. Formal definition
The paper defines a positional encoding as a function
mapping an -dimensional position to an embedding vector
where matches the token embedding dimension (Jabareen et al., 2 Sep 2025).
For reference, the paper gives the sinusoidal baseline as
$SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$
with temperature and . It gives isotropic Fourier Feature Positional Encoding as
where
is sampled from
0
The learnable Fourier baseline is
1
where
2
is trainable and initialized by
3
(Jabareen et al., 2 Sep 2025).
AFPE is obtained by replacing the single isotropic scale with per-dimension scales. The exact definition given is
4
where
5
and each 6 is the scale for the 7-th spatial dimension. This Gaussian is then used in the same Fourier mapping as IFPE: 8 The paper does not introduce a full covariance matrix or non-diagonal anisotropic Gaussian; its anisotropy model is strictly axis-wise scaling (Jabareen et al., 2 Sep 2025).
The construction is presented generically for 9-dimensional coordinates. For 2D images, 0; for 3D volumes, 1; and for spatiotemporal video, 2 with one temporal and two spatial coordinates. In all cases, the procedure is the same: sample or construct 3 according to per-axis scales, compute 4, concatenate sine and cosine features to obtain a 5-dimensional positional embedding, and add the embedding elementwise to patch tokens in a ViT pipeline (Jabareen et al., 2 Sep 2025).
3. Geometric interpretation and anisotropic inductive bias
The geometric point of AFPE is that not all coordinate directions should be treated equally. In IFPE, positional similarity is isotropic because one scale hyperparameter affects all axes equally. In AFPE, each axis has its own scale, so similarity can decay differently along different directions. The paper presents this as the ability to model stronger continuity along one axis, weaker continuity along another, and distinct spatial versus temporal similarity regimes in video (Jabareen et al., 2 Sep 2025).
The paper repeatedly connects this directional control to anatomical and pathological shape priors. It argues that elongated organs or patterns may benefit from anisotropic continuity, whereas round structures may favor isotropic settings. In the ChestX analysis, 6 works well for Atelectasis, Pneumothorax, Cardiomegaly, Tortuous Aorta, Subcutaneous Emphysema, and Pneumomediastinum; 7 works well for Emphysema, Pleural Thickening, Infiltration, Effusion, and Nodule; and 8 is reported as optimal for Mass, which the paper links to its round shape (Jabareen et al., 2 Sep 2025).
For EchoNet, the final selected values are
9
with the two spatial dimensions constrained to share the same scale. The paper interprets this as reflecting greater independence across time than across space (Jabareen et al., 2 Sep 2025).
A plausible interpretation is that AFPE replaces isotropic geometry with an axis-weighted geometry. The paper states that its definition is equivalent to a diagonal-covariance Gaussian over the projection vectors, but does not write the full multivariate distribution explicitly. This suggests that AFPE can be read as a stationary Fourier-feature encoding whose correlation structure differs by axis rather than remaining uniform in every direction (Jabareen et al., 2 Sep 2025).
This axis-sensitive view is related to learnable Fourier-feature formulations in which projection vectors define orientation and wavelength. “Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding” describes a trainable matrix 0 whose rows define both orientation and wavelength of Fourier features, but there anisotropy is emergent because 1 is unconstrained and learned rather than specified by explicit per-axis scales (Li et al., 2021).
4. Relation to adjacent positional encodings and common misconceptions
AFPE is best understood against three neighboring families: classical sinusoidal encodings, isotropic Fourier-feature encodings, and learnable Fourier-feature encodings. SPE encodes each axis separately and is criticized in the AFPE paper for struggling to preserve Euclidean distances in higher-dimensional spaces. IFPE improves multi-dimensional distance modeling, but assumes a single shared scale. LFPE makes the projection matrix trainable, but does not explicitly encode domain anisotropy through fixed axis-specific scales (Jabareen et al., 2 Sep 2025).
The general theoretical background comes from Fourier-feature work rather than from medical imaging alone. Tancik et al. show that Fourier feature mappings enable MLPs to learn high-frequency functions in low-dimensional domains, and characterize their effect through a stationary kernel with tunable bandwidth (Tancik et al., 2020). AFPE differs by shifting the design question from a single global bandwidth to a vector of per-dimension scales.
Several nearby arXiv papers are AFPE-relevant without actually proposing AFPE. “3QFP: Efficient neural implicit surface reconstruction using Tri-Quadtrees and Fourier feature Positional encoding” uses a standard Gaussian Fourier feature positional encoding with coefficients sampled from an isotropic Gaussian distribution; its Fourier encoding is explicitly isotropic, and any directional structure enters through tri-planar feature storage rather than through the Fourier map itself (Sun et al., 2024). “Rethinking Positional Encoding for Neural Vehicle Routing” proposes a hierarchical anisometric positional encoding built from distance-indexed sinusoidal in-route encoding and depot-anchored angular cross-route encoding; it is Fourier-like and anisometric, but it is not the medical-imaging AFPE formulation and does not use AFPE as an official term (Hua et al., 12 May 2026).
A common misconception is therefore to treat any directional or geometry-aware sinusoidal encoding as AFPE. In the strict sense established by (Jabareen et al., 2 Sep 2025), AFPE denotes an anisotropic generalization of IFPE in which dimension-specific Gaussian scales are used to sample Fourier projections for medical-image coordinates.
5. Empirical results
The paper evaluates AFPE on five tasks: multi-label classification on ChestX, organ classification on OrganMNIST3D, ejection-fraction regression on EchoNet Dynamic, and Feret-diameter regression on AdrenalMNIST3D and VesselMNIST3D. The compared positional encodings are None, Learnable, SPE, IFPE, LFPE, and AFPE (Jabareen et al., 2 Sep 2025).
The main quantitative results support a conditional claim rather than a universal one. In isotropic settings, AFPE is competitive but not always uniquely best; in anisotropic settings, AFPE degrades least and is typically best. The most pronounced improvement appears in spatiotemporal echocardiography, where temporal and spatial dimensions are explicitly decoupled (Jabareen et al., 2 Sep 2025).
| Task / setting | Metric | Key result |
|---|---|---|
| ChestX | AUPRC | AFPE 2; SPE 3; IFPE 4 |
| OrganMNIST3D, anisotropy 1 | AUROC | AFPE 5; SPE 6; IFPE 7 |
| OrganMNIST3D, anisotropy 4 | AUROC | AFPE 8; SPE 9; IFPE 0 |
| OrganMNIST3D, anisotropy 6 | AUROC | AFPE 1; SPE 2; IFPE 3 |
| OrganMNIST3D, anisotropy 8 | AUROC | AFPE 4; SPE 5; IFPE 6 |
| EchoNet, anisotropy 7 | 7 | AFPE 8; IFPE 9; SPE $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$0 |
The EchoNet result is the strongest single result in the paper. At anisotropy 7, AFPE reaches $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$1 in $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$2, compared with $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$3 for IFPE, $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$4 for SPE, $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$5 for LFPE, $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$6 for Learnable, and $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$7 for no positional encoding (Jabareen et al., 2 Sep 2025).
The Feret-diameter ablation is important because it separates generic shape encoding from anisotropy handling. On AdrenalMNIST3D and VesselMNIST3D, Fourier-based encodings massively outperform SPE for shape-descriptor learning, while IFPE and AFPE are broadly similar. For example, on AdrenalMNIST3D at anisotropy 3, SPE yields $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$8 for Feret minimum and $SPE_t(p) = \begin{cases} \sin(\frac{p}{t^{freq_j}) & \text{if } j \text{ mod } 2 = 0 \ \cos(\frac{p}{t^{freq_j}) & \text{else} \end{cases}$9 for Feret maximum, whereas IFPE yields 0 and 1, and AFPE yields 2 and 3. The paper interprets this as evidence that AFPE’s largest gains arise from anisotropy modeling rather than from a generic superiority in shape representation (Jabareen et al., 2 Sep 2025).
6. Implementation, limitations, and future directions
AFPE is introduced as a replacement for the positional encoding function rather than as an architectural change to the Transformer. It is added to ViT patch tokens in the usual way, so parameter count is effectively unchanged except for negligible positional-encoding specification, and training or inference complexity is unchanged relative to other fixed absolute positional encodings (Jabareen et al., 2 Sep 2025).
The method is mostly fixed and hyperparameterized rather than learned end-to-end. For ChestX and EchoNet, the paper reports 50 training runs with randomly sampled 4, each trained for 75 epochs, after which the best scales were selected on validation performance. For OrganMNIST3D, no hyperparameter optimization was done; instead,
5
The paper presents this as a direct practical heuristic when acquisition anisotropy is known (Jabareen et al., 2 Sep 2025).
The limitations are equally explicit. First, AFPE is hand-specified or tuned rather than learned automatically. Second, anisotropy is modeled only axis-wise: the formulation uses separate 6 per dimension but no full covariance matrix, rotation-aware anisotropy, or non-axis-aligned metric. Third, the benefits are strongest in anisotropic settings rather than universally dominant across all tasks. Fourth, the theoretical development is limited: the paper motivates AFPE geometrically and empirically but does not provide a full new kernel derivation or formal proof of anisotropic distance preservation (Jabareen et al., 2 Sep 2025).
The paper proposes future work on segmentation models such as SAM, object detection frameworks such as DETR, and end-to-end learning of scale parameters. It also states that the formulation itself is general even though the experiments are medical: beyond medical imaging, anisotropic coordinate structure also appears in videos, remote sensing, scientific simulation grids, and robotics spatiotemporal data. This suggests a broader relevance for AFPE, but the paper does not experimentally validate those domains (Jabareen et al., 2 Sep 2025).