Papers
Topics
Authors
Recent
Search
2000 character limit reached

Multiplicative Relative Position Embedding (M4M)

Updated 14 July 2026
  • M4M is a Transformer positional mechanism that replaces additive biases with a multiplicative tri-term interaction between content, query-relative, and key-relative positions.
  • It uses a learned relative embedding vector with clipped distances, enabling dynamic modulation of attention logits for richer token-to-token interactions.
  • Empirical evaluations show that M4M improves performance on benchmarks like SQuAD compared to both additive and scalar gating positional methods.

Searching arXiv for the core M4M paper and closely related positional encoding work. Multiplicative Relative Position Embedding (M4M) denotes a family of Transformer positional mechanisms in which relative position information enters the attention score through multiplicative interactions rather than through additive token-level position vectors or additive logit biases. The name is explicit in "Multiplicative Position-aware Transformer Models for Language Understanding" (Huang et al., 2021), where M4M is defined as the “M4 multiplicative” method. Closely related ideas appear earlier in "Improve Transformer Models with Better Relative Position Embeddings" (Huang et al., 2020), which did not use the term M4M but introduced relative position schemes in which positional factors multiplicatively gate or otherwise directly modulate query–key compatibility.

1. Origin and conceptual scope

Transformer self-attention is position agnostic unless positional information is supplied explicitly. In the standard formulation, for one head,

eij=(xiWQ)(xjWK)⊤dz,e_{ij} = \frac{(x_i W^Q)(x_j W^K)^\top}{\sqrt{d_z}},

and attention weights are obtained by a softmax over jj. Absolute position methods inject position at the input level, as in learned BERT-style embeddings xi=ti+wix_i = t_i + w_i, whereas relative methods define a pairwise term indexed by j−ij-i or a clipped variant of that displacement (Huang et al., 2020).

The immediate precursor of M4M is the relative-position line inaugurated by Shaw et al., where the attention score is

eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},

with aijKa_{ij}^K determined by a clipped signed distance (Shaw et al., 2018). This is additive at the logit level: content and position contribute as separate dot products. M4M departs from that decomposition by replacing additive accumulation with multiplicative interaction among content–content, query–position, and key–position compatibilities (Huang et al., 2021).

In the narrower sense used in the 2021 paper, M4M is a specific learned relative logit function. In a broader technical sense, the phrase also covers multiplicative relative position mechanisms such as scalar gating, tri-linear relative gating, and rotary schemes that encode relative offsets through multiplicative transformations of queries and keys (Huang et al., 2020, Su et al., 2021).

2. Formal definition inside attention logits

In the explicit M4M formulation, the model uses a learned relative embedding vector aij∈Rdza_{ij}\in\mathbb{R}^{d_z} indexed by clipped relative distance:

aij=wclip(j−i,k),clip(x,k)=max⁡(−k,min⁡(k,x)).a_{ij} = w_{\mathrm{clip}(j-i,k)}, \qquad \mathrm{clip}(x,k)=\max(-k,\min(k,x)).

Queries and keys are the standard linear projections

qi=xiWQ,kj=xjWK.q_i = x_i W^Q, \qquad k_j = x_j W^K.

M4M then defines the attention logit as (Huang et al., 2021)

eijM4M=(qi⋅kj)×(qi⋅aij)×(kj⋅aij)dz.e_{ij}^{\text{M4M}} = \frac{(q_i \cdot k_j)\times(q_i \cdot a_{ij})\times(k_j \cdot a_{ij})}{\sqrt{d_z}}.

This expression multiplies three scalar dot products:

  • jj0: content–content compatibility.
  • jj1: query–relative-position compatibility.
  • jj2: key–relative-position compatibility.

The value path is unchanged:

jj3

Accordingly, M4M is a positional mechanism in the attention logits rather than an additive embedding on the token representations (Huang et al., 2021).

The major positional alternatives discussed around M4M can be summarized as follows.

Method Logit form Interaction type
Absolute BERT/RoBERTa jj4 with position added in jj5 Additive at input
Shaw relative jj6 Additive in logits
Huang M4 jj7 Additive tri-term
M4M jj8 Multiplicative tri-term

This comparison makes the distinguishing feature precise: M4M does not merely attach an extra position bias to a content score; it makes the score proportional to the product of three compatibilities (Huang et al., 2021).

3. Precursors in relative position embedding research

Shaw et al. established the basic relative-position template by learning jj9 embeddings for signed distances in xi=ti+wix_i = t_i + w_i0, clipping larger distances into the boundary buckets. Their ablations showed that relative information injected into the logits, through xi=ti+wix_i = t_i + w_i1, was sufficient to obtain the full benefit on WMT 2014 English-to-German translation, and that combining relative and absolute positions yielded no further improvement (Shaw et al., 2018). That result fixed the importance of relative indexing and clipping as core design elements later inherited by M4M.

The 2020 paper on better relative position embeddings extends this line with four methods that are directly relevant to the later M4M nomenclature (Huang et al., 2020). Method 1 and Method 2 use a scalar xi=ti+wix_i = t_i + w_i2 to multiply the standard query–key dot product:

xi=ti+wix_i = t_i + w_i3

Method 1 uses unsigned distance xi=ti+wix_i = t_i + w_i4, while Method 2 uses signed distance xi=ti+wix_i = t_i + w_i5. These are pure multiplicative gates.

Method 3 upgrades the relative signal from a scalar to a vector and defines a tri-linear interaction

xi=ti+wix_i = t_i + w_i6

The paper states that the relative position embedding “serves as a gate to filter out the dot product of query and key” (Huang et al., 2020). This is an explicit multiplicative relative position mechanism in a dimensionwise form.

Method 4 is especially important because it combines vector-valued relative embeddings with three pairwise dot products:

xi=ti+wix_i = t_i + w_i7

Although this is additive rather than fully multiplicative, the same paper shows that it can be rewritten as

xi=ti+wix_i = t_i + w_i8

and argues that absolute position embeddings are a specific case of this construction (Huang et al., 2020). For that reason, Method 4 is often treated as the strongest practical precursor to M4M even though the fully multiplicative operator appears explicitly only in the later 2021 formulation.

4. Empirical performance and task profile

The 2021 M4M paper reports a systematic comparison on identically pretrained RoBERTa-base variants over MNLI-m, SST-2, and SQuAD1.1 dev (Huang et al., 2021). The absolute baseline obtains 81.25 on MNLI-m, 90.25 on SST-2, and 87.08 F1 on SQuAD1.1. Under the same setup, Shaw reaches 82.88, 91.28, and 88.62; M4 reaches 83.05, 91.05, and 89.36; DeBERTa reaches 83.67, 90.71, and 88.84; and M4M reaches 83.58, 91.62, and 89.45. In that table, M4M is the best SQuAD1.1 model and improves over M4 on all three tasks.

On larger-scale continued pretraining, RoBERTa-M4M is compared with RoBERTa and RoBERTa-ABS. On GLUE dev, the paper reports that RoBERTa-M4M reaches 87.82/87.59 on MNLI-m/mm, 92.98 on QNLI, 94.26 on SST-2, and 91.63 F1 on MRPC, while underperforming RoBERTa-ABS on CoLA and RTE (Huang et al., 2021). The task profile is therefore not uniformly dominant across all benchmarks.

The strongest gains are reported for extractive question answering. On SQuAD1.1 and SQuAD2.0, RoBERTa-M4M obtains 86.44 / 92.52 and 80.88 / 84.00, compared with 86.10 / 92.31 and 80.67 / 83.69 for RoBERTa-ABS. With larger C4-en pretraining, the paper reports 86.54 / 92.64 on SQuAD1.1 and 81.65 / 84.51 on SQuAD2.0 (Huang et al., 2021). On RoBERTa-large, M4M improves SQuAD1.1 from 94.63 F1 to 94.78 and SQuAD2.0 from 87.62 F1 to 88.34.

A recurring interpretation in the source literature is that M4M is particularly effective when token-to-token positional interactions matter directly. The 2020 precursor paper makes a closely related observation: on SQuAD1.1, vector-valued multiplicative or strongly coupled relative methods outperform both absolute embeddings and Shaw-style relative attention, while GLUE differences are smaller, plausibly because classification through the xi=ti+wix_i = t_i + w_i9 token reduces dependence on rich token-to-token positional structure (Huang et al., 2020).

5. Inductive properties, parameterization, and implementation

M4M inherits the clipped relative indexing scheme from Shaw et al. and therefore retains the same inductive bias: translation-invariant position encoding with exact distinction only inside a finite radius and coarse treatment beyond it (Shaw et al., 2018). In the 2020 paper, performance for a Shaw-style or Method 4 relative setup on SQuAD1.1 improves with larger clipping distance and saturates around j−ij-i0; for j−ij-i1, F1 is very similar, with examples including 90.30 at j−ij-i2, 90.54 at j−ij-i3, and 90.53 at j−ij-i4 (Huang et al., 2020). That result indicates that fine-grained long-distance distinctions contribute marginally beyond a moderate local window.

The same work also demonstrates an inductive advantage over absolute embeddings for longer sequences. A model pretrained with maximum length 512 and j−ij-i5 can be fine-tuned with maximum lengths 576, 640, and 704; absolute embeddings cannot do this because positions beyond 512 have no learned embeddings, whereas relative embeddings only require distances already covered by clipping (Huang et al., 2020).

In implementation terms, M4M is a drop-in replacement at the score computation. The 2021 paper keeps the Transformer block structure unchanged and only replaces the attention logit formula (Huang et al., 2021). Relative embedding tables are shared across heads within each layer in the reported implementation. For RoBERTa-base with j−ij-i6, j−ij-i7, j−ij-i8, and j−ij-i9, the paper gives the parameter count for M4M, M4, and Shaw as eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},0, which is negligible relative to the full model. The 2020 paper reports a similar conclusion for BERT-base, stating that Methods 3 and 4 add only about 147k parameters relative to BERT’s 108M parameters (Huang et al., 2020).

The literature also records an optimization contrast. In the RoBERTa experiments, additive scalar-bias variants such as Raffel and TUPE showed fine-tuning instability unless the learning rate was reduced, whereas the paper reports no convergence or stability problems with M4M (Huang et al., 2021).

6. Relation to rotary and group-theoretic multiplicative encodings

M4M belongs to a wider class of multiplicative relative position mechanisms in which position modulates attention through operators rather than additive biases. RoPE is the most prominent example of this broader class. In RoPE, position is encoded by multiplying queries and keys by block-diagonal 2D rotation matrices; the inner product then depends on positions only through the relative offset eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},1, because

eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},2

RoPE is therefore multiplicative and relative, but it is structurally different from M4M: it uses deterministic rotations in query/key space rather than learned clipped-distance vectors inside a tri-product logit (Su et al., 2021).

Later work generalizes the rotary idea in explicitly group-theoretic directions. LieRE maps positions eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},3 to skew-symmetric generators eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},4, exponentiates them to eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},5, and rotates queries and keys multiplicatively so that attention depends approximately on eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},6 through eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},7 (Ostmeier et al., 2024). Selective RoPE introduces input-dependent rotary transitions eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},8 and defines a relative composition eij=qi(kj+aijK)⊤dz=qikj⊤+qi(aijK)⊤dz,e_{ij} = \frac{q_i (k_j + a_{ij}^K)^\top}{\sqrt{d_z}} = \frac{q_i k_j^\top + q_i (a_{ij}^K)^\top}{\sqrt{d_z}},9, yielding

aijKa_{ij}^K0

which makes the multiplicative positional operator sequence dependent (Movahedi et al., 21 Nov 2025). GRAPE then frames multiplicative positional encoding as a group action in aijKa_{ij}^K1, with

aijKa_{ij}^K2

recovering RoPE exactly when the rotation planes are canonical coordinate pairs with log-uniform spectrum and extending it through learned commuting subspaces and compact non-commuting mixtures (Zhang et al., 8 Dec 2025).

These later developments clarify a common misconception. M4M is not synonymous with all relative position methods, and it is not identical to RoPE. Shaw-style attention is relative but additive at the logit level; RoPE is multiplicative and relative but matrix-rotational; the 2021 M4M method is a learned clipped-distance tri-product in the attention logits (Shaw et al., 2018, Su et al., 2021, Huang et al., 2021). What unifies these otherwise different mechanisms is the attempt to encode relative position in a form that acts directly on content compatibility rather than only as an external additive bias.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Multiplicative Relative Position Embedding (M4M).