---
title: Multiplicative Relative Position Embedding (M4M)
url: https://www.emergentmind.com/topics/multiplicative-relative-position-embedding-m4m
type: topic
---

# Multiplicative Relative Position Embedding (M4M)

Searching arXiv for the core M4M paper and closely related positional encoding work.
Multiplicative Relative Position Embedding (M4M) denotes a family of Transformer positional mechanisms in which relative position information enters the attention score through multiplicative interactions rather than through additive token-level position vectors or additive logit biases. The name is explicit in "Multiplicative Position-aware Transformer Models for Language Understanding" [2109.12788], where M4M is defined as the “M4 multiplicative” method. Closely related ideas appear earlier in "Improve Transformer Models with Better Relative Position Embeddings" [2009.13658], which did not use the term M4M but introduced relative position schemes in which positional factors multiplicatively gate or otherwise directly modulate query–key compatibility.

## 1. Origin and conceptual scope

Transformer self-attention is position agnostic unless positional information is supplied explicitly. In the standard formulation, for one head,
$$
e_{ij} = \frac{(x_i W^Q)(x_j W^K)^\top}{\sqrt{d_z}},
$$
and attention weights are obtained by a softmax over \(j\). Absolute position methods inject position at the input level, as in learned BERT-style embeddings \(x_i = t_i + w_i\), whereas relative methods define a pairwise term indexed by \(j-i\) or a clipped variant of that displacement [2009.13658].

The immediate precursor of M4M is the relative-position line inaugurated by Shaw et al., where the attention score is
$$
e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}}
      = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},
$$
with \(a_{ij}^K\) determined by a clipped signed distance [1803.02155]. This is additive at the logit level: content and position contribute as separate dot products. M4M departs from that decomposition by replacing additive accumulation with multiplicative interaction among content–content, query–position, and key–position compatibilities [2109.12788].

In the narrower sense used in the 2021 paper, M4M is a specific learned relative logit function. In a broader technical sense, the phrase also covers multiplicative relative position mechanisms such as scalar gating, tri-linear relative gating, and rotary schemes that encode relative offsets through multiplicative transformations of queries and keys [2009.13658; 2104.09864].

## 2. Formal definition inside attention logits

In the explicit M4M formulation, the model uses a learned relative embedding vector \(a_{ij}\in\mathbb{R}^{d_z}\) indexed by clipped relative distance:
$$
a_{ij} = w_{\mathrm{clip}(j-i,k)}, \qquad \mathrm{clip}(x,k)=\max(-k,\min(k,x)).
$$
Queries and keys are the standard linear projections
$$
q_i = x_i W^Q, \qquad k_j = x_j W^K.
$$
M4M then defines the attention logit as [2109.12788]
$$
e_{ij}^{\text{M4M}}
=
\frac{(q_i \cdot k_j)\times(q_i \cdot a_{ij})\times(k_j \cdot a_{ij})}{\sqrt{d_z}}.
$$

This expression multiplies three scalar dot products:

- \(q_i \cdot k_j\): content–content compatibility.
- \(q_i \cdot a_{ij}\): query–relative-position compatibility.
- \(k_j \cdot a_{ij}\): key–relative-position compatibility.

The value path is unchanged:
$$
z_i = \sum_j \alpha_{ij}(x_j W^V), \qquad \alpha_{ij}=\mathrm{softmax}_j(e_{ij}^{\text{M4M}}).
$$
Accordingly, M4M is a positional mechanism in the attention logits rather than an additive embedding on the token representations [2109.12788].

The major positional alternatives discussed around M4M can be summarized as follows.

| Method | Logit form | Interaction type |
|---|---|---|
| Absolute BERT/RoBERTa | \(\frac{(x_iW^Q)(x_jW^K)^\top}{\sqrt{d_z}}\) with position added in \(x_i\) | Additive at input |
| Shaw relative | \(\frac{q_i\cdot(k_j+a_{ij})}{\sqrt{d_z}}\) | Additive in logits |
| Huang M4 | \(\frac{q_i\cdot k_j + q_i\cdot a_{ij} + k_j\cdot a_{ij}}{\sqrt{d_z}}\) | Additive tri-term |
| M4M | \(\frac{(q_i\cdot k_j)(q_i\cdot a_{ij})(k_j\cdot a_{ij})}{\sqrt{d_z}}\) | Multiplicative tri-term |

This comparison makes the distinguishing feature precise: M4M does not merely attach an extra position bias to a content score; it makes the score proportional to the product of three compatibilities [2109.12788].

## 3. Precursors in relative position embedding research

Shaw et al. established the basic relative-position template by learning \(2k+1\) embeddings for signed distances in \([-k,k]\), clipping larger distances into the boundary buckets. Their ablations showed that relative information injected into the logits, through \(a_{ij}^K\), was sufficient to obtain the full benefit on WMT 2014 English-to-German translation, and that combining relative and absolute positions yielded no further improvement [1803.02155]. That result fixed the importance of relative indexing and clipping as core design elements later inherited by M4M.

The 2020 paper on better relative position embeddings extends this line with four methods that are directly relevant to the later M4M nomenclature [2009.13658]. Method 1 and Method 2 use a scalar \(a_{ij}\) to multiply the standard query–key dot product:
$$
e_{ij} = \frac{(q_i k_j^\top)a_{ij}}{\sqrt{d_z}}.
$$
Method 1 uses unsigned distance \(a_{ij}=w_{|j-i|}\), while Method 2 uses signed distance \(a_{ij}=w_{j-i}\). These are pure multiplicative gates.

Method 3 upgrades the relative signal from a scalar to a vector and defines a tri-linear interaction
$$
e_{ij} = \frac{\mathrm{sum\_prod}(q_i, k_j, a_{ij})}{\sqrt{d_z}}
       = \frac{\sum_{d=1}^{d_z} q_{i,d}k_{j,d}a_{ij,d}}{\sqrt{d_z}}.
$$
The paper states that the relative position embedding “serves as a gate to filter out the dot product of query and key” [2009.13658]. This is an explicit multiplicative relative position mechanism in a dimensionwise form.

Method 4 is especially important because it combines vector-valued relative embeddings with three pairwise dot products:
$$
e_{ij}
=
\frac{(x_iW^Q)\cdot(x_jW^K) + (x_iW^Q)\cdot a_{ij} + (x_jW^K)\cdot a_{ij}}{\sqrt{d_z}}.
$$
Although this is additive rather than fully multiplicative, the same paper shows that it can be rewritten as
$$
e_{ij}
=
\frac{(x_iW^Q + a_{ij})(x_jW^K + a_{ij})^\top - a_{ij}a_{ij}^\top}{\sqrt{d_z}},
$$
and argues that absolute position embeddings are a specific case of this construction [2009.13658]. For that reason, Method 4 is often treated as the strongest practical precursor to M4M even though the fully multiplicative operator appears explicitly only in the later 2021 formulation.

## 4. Empirical performance and task profile

The 2021 M4M paper reports a systematic comparison on identically pretrained RoBERTa-base variants over MNLI-m, SST-2, and SQuAD1.1 dev [2109.12788]. The absolute baseline obtains 81.25 on MNLI-m, 90.25 on SST-2, and 87.08 F1 on SQuAD1.1. Under the same setup, Shaw reaches 82.88, 91.28, and 88.62; M4 reaches 83.05, 91.05, and 89.36; DeBERTa reaches 83.67, 90.71, and 88.84; and M4M reaches 83.58, 91.62, and 89.45. In that table, M4M is the best SQuAD1.1 model and improves over M4 on all three tasks.

On larger-scale continued pretraining, RoBERTa-M4M is compared with RoBERTa and RoBERTa-ABS. On GLUE dev, the paper reports that RoBERTa-M4M reaches 87.82/87.59 on MNLI-m/mm, 92.98 on QNLI, 94.26 on SST-2, and 91.63 F1 on MRPC, while underperforming RoBERTa-ABS on CoLA and RTE [2109.12788]. The task profile is therefore not uniformly dominant across all benchmarks.

The strongest gains are reported for extractive question answering. On SQuAD1.1 and SQuAD2.0, RoBERTa-M4M obtains 86.44 / 92.52 and 80.88 / 84.00, compared with 86.10 / 92.31 and 80.67 / 83.69 for RoBERTa-ABS. With larger C4-en pretraining, the paper reports 86.54 / 92.64 on SQuAD1.1 and 81.65 / 84.51 on SQuAD2.0 [2109.12788]. On RoBERTa-large, M4M improves SQuAD1.1 from 94.63 F1 to 94.78 and SQuAD2.0 from 87.62 F1 to 88.34.

A recurring interpretation in the source literature is that M4M is particularly effective when token-to-token positional interactions matter directly. The 2020 precursor paper makes a closely related observation: on SQuAD1.1, vector-valued multiplicative or strongly coupled relative methods outperform both absolute embeddings and Shaw-style relative attention, while GLUE differences are smaller, plausibly because classification through the \([CLS]\) token reduces dependence on rich token-to-token positional structure [2009.13658].

## 5. Inductive properties, parameterization, and implementation

M4M inherits the clipped relative indexing scheme from Shaw et al. and therefore retains the same inductive bias: translation-invariant position encoding with exact distinction only inside a finite radius and coarse treatment beyond it [1803.02155]. In the 2020 paper, performance for a Shaw-style or Method 4 relative setup on SQuAD1.1 improves with larger clipping distance and saturates around \(k \approx 32\); for \(k \ge 32\), F1 is very similar, with examples including 90.30 at \(k=32\), 90.54 at \(k=256\), and 90.53 at \(k=512\) [2009.13658]. That result indicates that fine-grained long-distance distinctions contribute marginally beyond a moderate local window.

The same work also demonstrates an inductive advantage over absolute embeddings for longer sequences. A model pretrained with maximum length 512 and \(k=256\) can be fine-tuned with maximum lengths 576, 640, and 704; absolute embeddings cannot do this because positions beyond 512 have no learned embeddings, whereas relative embeddings only require distances already covered by clipping [2009.13658].

In implementation terms, M4M is a drop-in replacement at the score computation. The 2021 paper keeps the Transformer block structure unchanged and only replaces the attention logit formula [2109.12788]. Relative embedding tables are shared across heads within each layer in the reported implementation. For RoBERTa-base with \(m=12\), \(n=512\), \(d=768\), and \(h=12\), the paper gives the parameter count for M4M, M4, and Shaw as \(m(2n-1)d/h \approx 785\text{K}\), which is negligible relative to the full model. The 2020 paper reports a similar conclusion for BERT-base, stating that Methods 3 and 4 add only about 147k parameters relative to BERT’s 108M parameters [2009.13658].

The literature also records an optimization contrast. In the RoBERTa experiments, additive scalar-bias variants such as Raffel and TUPE showed fine-tuning instability unless the learning rate was reduced, whereas the paper reports no convergence or stability problems with M4M [2109.12788].

## 6. Relation to rotary and group-theoretic multiplicative encodings

M4M belongs to a wider class of multiplicative relative position mechanisms in which position modulates attention through operators rather than additive biases. RoPE is the most prominent example of this broader class. In RoPE, position is encoded by multiplying queries and keys by block-diagonal 2D rotation matrices; the inner product then depends on positions only through the relative offset \(n-m\), because
$$
(R_{\Theta,m}^d W_q x_m)^\top(R_{\Theta,n}^d W_k x_n)
=
(W_q x_m)^\top R_{\Theta,n-m}^d (W_k x_n).
$$
RoPE is therefore multiplicative and relative, but it is structurally different from M4M: it uses deterministic rotations in query/key space rather than learned clipped-distance vectors inside a tri-product logit [2104.09864].

Later work generalizes the rotary idea in explicitly group-theoretic directions. LieRE maps positions \(x\in\mathbb{R}^d\) to skew-symmetric generators \(P(x)=Ax\), exponentiates them to \(R(x)=\exp(Ax)\in SO(n)\), and rotates queries and keys multiplicatively so that attention depends approximately on \(x_i-x_j\) through \(R(x_j)^{-1}R(x_i)\) [2406.10322]. Selective RoPE introduces input-dependent rotary transitions \(R_t\) and defines a relative composition \(R_{\tau+1:t}=\prod_{\kappa=\tau+1}^t R_\kappa\), yielding
$$
o_t = \sum_{\tau=1}^{t} v_\tau\,\{k_\tau^\top R_{\tau+1:t} q_t\},
$$
which makes the multiplicative positional operator sequence dependent [2511.17388]. GRAPE then frames multiplicative positional encoding as a group action in \(SO(d)\), with
$$
\mathbf{G}(n)=\exp(n\,\omega\,\mathbf{L}),
$$
recovering RoPE exactly when the rotation planes are canonical coordinate pairs with log-uniform spectrum and extending it through learned commuting subspaces and compact non-commuting mixtures [2512.07805].

These later developments clarify a common misconception. M4M is not synonymous with all relative position methods, and it is not identical to RoPE. Shaw-style attention is relative but additive at the logit level; RoPE is multiplicative and relative but matrix-rotational; the 2021 M4M method is a learned clipped-distance tri-product in the attention logits [1803.02155; 2104.09864; 2109.12788]. What unifies these otherwise different mechanisms is the attempt to encode relative position in a form that acts directly on content compatibility rather than only as an external additive bias.

Source: https://www.emergentmind.com/topics/multiplicative-relative-position-embedding-m4m