Papers
Topics
Authors
Recent
Search
2000 character limit reached

Three-Way Weight Tying (TWWT) in NMT

Updated 27 January 2026
  • Three-Way Weight Tying (TWWT) is a parameter-sharing strategy that integrates input embeddings, output classifiers, and decoder-context projections into a joint embedding space.
  • It enhances translation performance, especially for morphologically rich languages, by enabling flexible control over the output-layer capacity.
  • Empirical evaluations demonstrate consistent BLEU improvements over baseline models while ensuring robustness across different network depths and vocabulary sizes.

Three-way weight tying (TWWT) is a parameter-sharing strategy introduced for neural machine translation (NMT) models, specifically enhancing the standard attention-based encoder–decoder architecture. TWWT generalizes and extends conventional weight tying—where input embeddings and output classifiers share parameters—by introducing a joint input-output embedding. This mechanism not only ties input embeddings and output classifiers, but also the decoder-context projection, into a shared parameter space. The resulting structure-aware output layer enables explicit control over model capacity and demonstrates superior empirical performance in translation tasks, notably for morphologically rich languages (Pappas et al., 2018).

1. Standard Attention-Based NMT and Weight Tying

In the baseline NMT setup, the encoder transforms source word indices x1…xmx_1\dots x_m into embedding vectors via E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}, producing encoder hidden states h1e,…,hmeh^e_1, \dots, h^e_m using LSTMs or bi-LSTMs. At each decoding step tt, the decoder LSTM state ht∈Rdhh_t \in \mathbb{R}^{d_h} is computed using the embedding of the previous target token and the prior attention context ct−1c_{t-1}. The attention mechanism computes context vectors as weighted sums over encoder states, with weights determined by

αti=softmaxi(ht⊤Wahie)\alpha_{t i} = \mathrm{softmax}_i(h_t^\top W_a h^e_i)

and

ct=∑iαtihie.c_t = \sum_i \alpha_{t i} h^e_i.

The decoder output hth_t is projected through a softmax-linear layer,

p(yt∣y1:t−1,X)∝exp⁡(W⊤ht+b),p(y_t|y_{1:t-1},X) \propto \exp(W^\top h_t + b),

with E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}0 and E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}1. In conventional weight tying, E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}2 enforces equality between input embeddings and output classifiers, improving efficiency and often translation quality.

2. Joint Input–Output Embedding Formulation

TWWT replaces the linear output projection with a nonlinear joint embedding approach. Each target word embedding and decoder hidden state is nonlinearly projected into a shared E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}3-dimensional joint space. Specifically, for embedding E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}4 and decoder state E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}5,

E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}6

E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}7

with E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}8, E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d}9, h1e,…,hmeh^e_1, \dots, h^e_m0, and h1e,…,hmeh^e_1, \dots, h^e_m1; h1e,…,hmeh^e_1, \dots, h^e_m2 is a nonlinearity (tanh in empirical studies). The score for candidate h1e,…,hmeh^e_1, \dots, h^e_m3 at position h1e,…,hmeh^e_1, \dots, h^e_m4 is

h1e,…,hmeh^e_1, \dots, h^e_m5

where h1e,…,hmeh^e_1, \dots, h^e_m6. Thus, prediction probabilities become

h1e,…,hmeh^e_1, \dots, h^e_m7

In matrix notation, letting h1e,…,hmeh^e_1, \dots, h^e_m8 and h1e,…,hmeh^e_1, \dots, h^e_m9, the output is tt0.

3. Parameter Tying in TWWT

TWWT introduces a three-way tying among:

  • The input embedding parameters tt1
  • The output classifier projection tt2
  • The decoder-context projection tt3

These can be tied to a single shared matrix tt4 and bias tt5: tt6 yielding

tt7

This structured parameter-sharing governs not only input and output mappings but the transformation of the decoder's context. Additionally, optional residual or gating components can be added to relax the hard tie, for example: tt8 with learned gates tt9.

4. Control of Output-Layer Capacity

TWWT allows flexible capacity adjustment through the joint space dimensionality ht∈Rdhh_t \in \mathbb{R}^{d_h}0, interpolating from low-capacity, tightly regularized models to high-capacity models akin to an unrestricted softmax layer. The number of parameters in the joint output layer is

ht∈Rdhh_t \in \mathbb{R}^{d_h}1

where ht∈Rdhh_t \in \mathbb{R}^{d_h}2 is the vocabulary size. Varying ht∈Rdhh_t \in \mathbb{R}^{d_h}3 thus trades off between model compactness and expressivity, with the regimes ordered as ht∈Rdhh_t \in \mathbb{R}^{d_h}4. Notably, adjusting ht∈Rdhh_t \in \mathbb{R}^{d_h}5 does not require re-architecting the overall network.

5. Empirical Evaluation and Implementation

Experiments evaluate TWWT in English–Finnish and English–German translation. Empirical setup includes:

  • Vocabulary ht∈Rdhh_t \in \mathbb{R}^{d_h}6 via BPE.
  • Embedding size ht∈Rdhh_t \in \mathbb{R}^{d_h}7, decoder LSTM hidden size ht∈Rdhh_t \in \mathbb{R}^{d_h}8.
  • Joint space dimensions ht∈Rdhh_t \in \mathbb{R}^{d_h}9, selected on development sets.
  • Stacked LSTMs: 2 layers baseline, with additional experiments at 1, 4, and 8 layers.
  • Dropout ct−1c_{t-1}0 after LSTM layers, Adam optimizer (ct−1c_{t-1}1), negative sampling on large vocabularies (25% for ct−1c_{t-1}2, 75% for ct−1c_{t-1}3).

Translation quality (BLEU) improvements are summarized as follows. For ct−1c_{t-1}4 BPE:

Task Baseline Weight-Tied Joint/TWWT
English→Finnish 12.68 12.58 13.03*
Finnish→English 9.42 9.59 10.19*
English→German 18.46 18.48 19.79*
German→English 15.85 16.51* 18.11***

(* ct−1c_{t-1}5, ** ct−1c_{t-1}6). BLEU gains with TWWT consistently exceed those of conventional tying and baseline models, with up to 2 BLEU improvement in morphologically rich languages. Increasing vocabulary size maintains TWWT’s empirical benefit (up to 1 BLEU). Training throughput (ct−1c_{t-1}7–ct−1c_{t-1}8k tokens/sec for ct−1c_{t-1}9) is comparable to baselines, and higher αti=softmaxi(ht⊤Wahie)\alpha_{t i} = \mathrm{softmax}_i(h_t^\top W_a h^e_i)0 introduces only modest slowdowns, largely mitigated by negative sampling. Depth and frequency robustness are also superior.

6. Significance and Model Implications

TWWT ensures that parameter sharing captures richer semantic structure and preserves prior knowledge in both input representations and translation contexts. The explicit low-rank nonlinear joint space facilitates smooth interpolation between strongly regularized and fully expressive output layers, without architectural disruption. This not only leads to superior performance on several language pairs but also offers robustness across network depths and vocabulary settings. The approach demonstrates that three-way parameter sharing—across embeddings, output classifiers, and context projections—provides a powerful regularization tool and improves translation quality and model robustness, particularly for morphologically complex tasks (Pappas et al., 2018).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Three-Way Weight Tying (TWWT).