---
title: 2D Rotary Position Embeddings (RoPE)
url: https://www.emergentmind.com/topics/2d-rotary-position-embeddings-rope
type: topic
---

# 2D Rotary Position Embeddings (RoPE)

Two-dimensional Rotary Position Embeddings (2D RoPE) generalize the foundational concept of rotary position encoding from 1D sequences to spatial or higher-dimensional domains, enabling Transformer-based architectures to encode continuous, relative, multi-axis position information directly and efficiently into their attention mechanisms. Originating from the block-diagonal rotation formulation introduced in RoFormer, recent advances have both formalized and substantially extended the expressive capacity, robustness, and geometric fidelity of RoPE for image, multimodal, and geometric reasoning tasks.

## 1. Mathematical Foundations and General RoPE Equation

2D RoPE seeks a position-dependent transformation $R(\mathbf{x}) \in \mathrm{SO}(d)$ such that for all positions $\mathbf{x}, \mathbf{y} \in \mathbb{R}^2$:
$$
R(\mathbf{x})^\top R(\mathbf{y}) = R(\mathbf{y} - \mathbf{x})
$$
This “RoPE Equation” ensures that the dot product between position-encoded queries and keys in self-attention depends only on the relative offset, yielding translation-invariant logits and seamless extrapolation to unseen resolution or field-of-view. The most general solution is to construct $R(\mathbf{x})$ via the matrix exponential of a linear map generated by pairwise commuting skew-symmetric matrices $\{A_i\}$ acting on each coordinate axis:
$$
R(\mathbf{x}) = \exp \left( \sum_{i=1}^N A_i x_i \right)
$$
with $[A_i, A_j] = 0$ for all $i, j$ [2506.03737, 2502.02562, 2504.06308].

In the block-diagonal case (maximal toral subalgebra), each $A_i$ acts on a disjoint $2 \times 2$ or $b \times b$ block, and $R(\mathbf{x})$ factors as independent rotations per axis, corresponding to standard axis-aligned 2D RoPE [2104.09864, 2403.13298, 2502.02562]. If the commutativity requirement is violated, as in generic learned Lie rotations, relative position dependence is lost and robustness deteriorates [2506.03737].

## 2. Constructive Parameterizations and Model Variants

Several parameterizations have been developed to instantiate or generalize 2D RoPE while maintaining the critical commutativity condition:

- **Axis-Aligned RoPE (Block-Diagonal):** Each input feature is split into halves, with independent $1$D RoPE applied on $x$ and $y$ axes using block-diagonal $2 \times 2$ rotations at log-uniform frequencies. This construction is computationally efficient, parameter-free, and translation-invariant, but cannot express diagonal or off-axis interactions [2104.09864, 2403.13298].

- **ComRoPE (“Commuting RoPE”):** Generalizes 2D RoPE by introducing trainable block-diagonal angle matrices $A_1, A_2$ for each coordinate, with enforced pairwise commutativity. ComRoPE-AP (axial partition) uses mutually exclusive block supports; ComRoPE-LD (linearly dependent) defines $A_i = \theta_i M$ for a single base generator $M$. Both maintain relative-position robustness and empirical performance beyond fixed-frequency RoPE [2506.03737].

- **STRING (Universal Lie Exponential):** Extends RoPE to arbitrary coordinate dimensionality by exponentiating a sum of $N$ commuting skew-symmetric generators, providing a universal construction for all translation-invariant, separable position encodings. This formalism unifies axial, diagonal, and block-structured RoPE as special cases [2502.02562].

- **Learned Diagonal/LieRE:** LieRE introduces a fully learned linear map from displacement vectors to the Lie algebra so(2), allowing free scalar weights but sacrificing the commutativity guarantee. This increases flexibility but can degrade large-offset generalization in higher-dimensional domains or under random coordinate perturbations [2406.10322, 2506.03737].

- **Spiral RoPE:** Overcomes axis-aligned limitations by partitioning embedding channels into multiple directional groups, each rotated according to the patch’s projection onto a uniformly distributed set of directions. This approach covers the entire frequency plane and better encodes oblique and curved spatial relationships without added parameters or computational cost [2602.03227].

- **GeoPE (Quaternionic RoPE):** Lifts 2D positions to 3D using quaternionic representations and constructs a symmetric (commuting in Lie algebra) joint rotation by averaging log-maps of axis-rotations, thus capturing true 2D spatial topology and shape bias. This avoids the “false neighbor” problem induced by flattening [2512.04963].

## 3. Implementation and Computational Complexity

The practical implementation of 2D RoPE retains the linear complexity and negligible memory overhead characteristic of the 1D variant, with computational cost dominated by the final matrix multiplication in attention:
- **Block-diagonal axis-aligned and ComRoPE**: Per-token computational overhead is $O(d)$ with $O(n \cdot d)$ total extra time per layer ($n$ tokens, $d$ dimension). Additional parameters per layer are $O(d \cdot b)$ (ComRoPE-AP) or $O(d (b + 2/b))$ (ComRoPE-LD), both modest compared to vanilla transformer parameter budgets [2506.03737].
- **STRING, LieRE:** Arbitrarily complex generator families require at worst $O(n d^3)$ (naïve), but efficient reductions (sparse/FFT/Cayley basis) achieve $O(n d)$ or $O(n d \log d)$ [2502.02562, 2504.06308].
- **Spiral RoPE:** Parameter-free and FLOPs-equivalent to block-diagonal RoPE [2602.03227].

No variant above materially increases attention memory footprint relative to absolute position embeddings or O($n^2 d$) RPE tables.

## 4. Applications and Empirical Performance

2D RoPE and its generalizations have achieved state-of-the-art or superior performance across a spectrum of modalities and tasks, particularly where translation-invariance, spatial extrapolation, and multi-scale structure are essential. Key empirical findings include:

- **Image Classification (ImageNet-1K):** ComRoPE-LD achieves 65.49% top-1 at 224x224 and 55.29% at 512x512, outperforming standard RoPE (+2.4% at train, +4.2% at extrapolated resolution) and LieRE (+1.6% train, +2.9% high-res) [2506.03737]. Spiral RoPE yields consistent gains up to +0.88% top-1 (ViT-B, 384x384) over axis-aligned, and semantically crisper attention [2602.03227].

- **Semantic Segmentation and Object Detection:** Spiral RoPE (+2.21% mIoU at 512x512 compared to APE on ADE20k) and GeoPE (+0.3%–1.1% absolute) show measurable, robust improvements [2602.03227, 2512.04963].

- **Vision-Language and Multimodal Tasks:** Both STRING-based and quaternion-based encodings enhance recall/IoU, especially in geometric and retrieval scenarios [2502.02562, 2512.04963].

- **Trajectory and Agent-Centric Modeling:** Directional RoPE (DRoPE) for agent heading breaks the periodicity/ambiguity of 1D RoPE and achieves state-of-the-art accuracy/efficiency trade-off in autonomous driving benchmarks, without quadratic space overhead [2503.15029].

## 5. Geometric and Theoretical Properties

A foundational property of 2D RoPE is its strict dependence of attention scores on *relative* position, guaranteed by the mathematical structure: $R(\mathbf{x})^\top R(\mathbf{y}) = R(\mathbf{y} - \mathbf{x})$. The underlying requirements are:
- **Commutativity of Generators:** Required for scalability to high-dimensional or multidimensional input; encodes the “translation invariance” at the heart of robust, resolution-agnostic vision transformers [2506.03737, 2502.02562].
- **Embedding in Maximal Abelian Subalgebras (MASA) of $\mathfrak{so}(d)$:** All valid 2D RoPEs correspond to choosing bases in these subalgebras; block-diagonal (axis-aligned) forms are the maximal toral case [2504.06308].
- **Norm Preservation, Compositionality:** All RoPEs (when constructed via orthogonal exponentials) preserve vector norms and admit streaming/caching by virtue of their group structure [2512.07805].
- **Diagonal versus Mixed-Directionality:** Axis-aligned RoPE is limited to frequencies lying on principal axes; Spiral RoPE, GeoPE, and diagonal-mixed variants provide richer, isotropic frequency coverage and better boundary, shape, and objectness encodings [2602.03227, 2512.04963].

## 6. Extensions and Recent Directions

2D RoPE has seen multiple generalizations and practical adaptations:

- **Cross-View and Cross-Dimensional Extensions (URoPE):** Universal RoPE lifts features to 3D with camera-depth anchors and projects back to 2D for cross-view transformers. This construction is parameter-free, SE(3)-invariant, and compatible with RoPE-optimized kernels [2604.18747].

- **Time-and-Order RoPE (TO-RoPE):** For generative recommendation, TO-RoPE jointly encodes ordinal sequence index and wall-clock time via rotary angles, supporting early fusion and split-by-dimension strategies. Empirically, this increases retrieval quality and broadens attention span over both axes [2510.20455].

- **Quaternionic and Non-Commutative Variants:** GeoPE fuses SO(3) rotations via Lie algebra averaging, ensuring symmetric treatment of spatial axes and outperforming axial approaches in tasks requiring true geometric awareness [2512.04963].

- **N-Dimensional and Orthogonal-Mixing Frameworks:** Complete theoretical unification places all RoPE variants in the context of MASA and Lie-algebra topology, enabling learned basis changes (e.g., via the Cayley transform) for greater coordination between position axes without breaking the essential commutativity/relativity guarantees [2504.06308].

## 7. Comparative Overview and Practical Integration

A comparative summary of several prominent 2D RoPE architectures is presented below:

| Method         | Generator Structure       | Relative Law | Param Count      | Empirical Gain            |
|----------------|--------------------------|--------------|------------------|---------------------------|
| Axis-Aligned   | Block-diag commuting     | Yes          | 0                | Baseline; robust, limited|
| Mixed/Flexible | Learned diagonal/mixed   | Yes/Partial  | $O(d)$           | +1–2% on ViT, multi-res  |
| ComRoPE        | Trainable commuting      | Yes          | $O(d \cdot b)$   | SOTA; most robust        |
| LieRE          | Arbitrary SO(2) learning | No           | $O(d)$           | Flexible, can degrade    |
| Spiral RoPE    | Multi-directional splits | Yes          | 0                | +0.54–0.88% (ImageNet)   |
| GeoPE          | SO(3) quaternionic mean  | Yes (linear) | 0                | +0.3–1.1% (ViT, Swin)    |
| DRoPE          | Uniform angular block    | Yes (angle)  | 0                | SOTA for agent heading   |

In practical deployment, 2D RoPE variants are typically integrated immediately after the $Q, K$ projections at each self-attention layer, with configured frequencies/axes and commutativity constraints enforced at model initialization.

---

Two-dimensional Rotary Position Embeddings, through rigorous Lie-theoretic foundations and empirical validation across scales and domains, have become the primary standard for robust, scalable, translation-invariant positional encoding in transformer models for vision, geometric, and multi-modal applications, with commutativity, block-diagonal structure, and relative law preservation as the defining technical features [2506.03737, 2502.02562, 2602.03227, 2504.06308, 2512.04963].

Source: https://www.emergentmind.com/topics/2d-rotary-position-embeddings-rope