---
title: 'MHCADDI: Multi-Head Co-Attentive DDI Encoder'
url: https://www.emergentmind.com/topics/multi-head-co-attentive-drug-drug-interaction-encoder-mhcaddi
type: topic
---

# MHCADDI: Multi-Head Co-Attentive DDI Encoder

A Multi-Head Co-Attentive Drug-Drug Interaction Encoder (MHCADDI) is a neural network architecture specifically designed to predict adverse effects arising from drug–drug interactions (DDIs) by operating directly on molecular graph representations of drug pairs. It leverages a novel integration of message-passing neural networks (MPNNs) with multi-head cross-drug co-attention mechanisms. The principal innovation is the incorporation of joint information from both drugs in the pair as early as possible when constructing atom-level representations, allowing for finer modeling of molecular interactions that may result in side effects [1905.00534].

## 1. Molecular Graph Representation

Each drug $d_x$ is modeled as an undirected molecular graph. The atomic structure is encoded as follows:

- **Node (Atom) Features**: Each atom $a_i^{(d_x)}$ includes:
    1. Atom type identifier (one-hot, projected to 32 dimensions)
    2. Number of bonded hydrogen atoms
    3. Formal atomic charge
    4. A learned 32-dimensional embedding per atom type

- **Edge (Bond) Features**: Each bond $e_{ij}^{(d_x)}$ is represented by a learned 32-dimensional embedding of bond category (single, double, etc.)

- **Preprocessing Pipeline**:
    1. Retrieve molecular graph via PubChem ID and extract atom/bond attributes
    2. Negative sampling strategy: In binary classification, generate negative triplets by corrupting one drug in each positive $(d_x, d_y, se_z)$ trio, where $se_z$ denotes the side-effect label.

## 2. Message Passing and Co-Attention

The MHCADDI architecture alternates between intra-drug message-passing and cross-drug co-attention for $T=3$ layers.

### 2.1. Intra-Drug Message Passing

Each atom feature is initialized by a projection: ${}^{(d_x)}h_i^{0} = f_i(a_i^{(d_x)}) \in \mathbb{R}^{32}$.

At each step $t$, messages are computed as:
\[
{}^{(d_x)}m_{ij}^t = f_e^t(e_{ij}^{(d_x)}) \odot f_v^t({}^{(d_x)}h_j^{t-1})
\]
where $f_e^t$ is a two-layer LeakyReLU MLP (each $32 \to 32$), $f_v^t$ is a single-layer projection ($32 \to 32$). Aggregation is by summation over neighbors.

### 2.2. Cross-Drug Co-Attention

For each atom $i$ in $d_x$ and $j$ in $d_y$:

Per head $k$:
\[
\begin{aligned}
{}^{(k)}q_i^t &= {}^{(k)}W_k^t\,{}^{(d_x)}h_i^{t-1}\\
{}^{(k)}k_j^t &= {}^{(k)}W_k^t\,{}^{(d_y)}h_j^{t-1}\\
{}^{(k)}v_j^t &= {}^{(k)}W_v^t\,{}^{(d_y)}h_j^{t-1}
\end{aligned}
\]
Attention weights are:
\[
{}^{(k)}\alpha_{ij}^t = \mathrm{softmax}_j(\langle {}^{(k)}q_i^t, {}^{(k)}k_j^t \rangle)
\]
The attended message is:
\[
{}^{(k)}n_{i}^t = \sum_{j \in d_y} {}^{(k)}\alpha_{ij}^t\,{}^{(k)}v_j^t
\]
Outputs from the $K=8$ heads are concatenated to a $256$-dimensional vector and linearly projected (via $f_o^t$) back to $32$ dimensions for each atom.

An identical block is applied in the reverse direction ($d_y \leftarrow d_x$).

### 2.3. Feature Update

The atom feature update includes normalization and residuals:
\[
{}^{(d_x)}h_i^t = \mathrm{LayerNorm}\left({}^{(d_x)}h_i^{t-1} + {}^{(d_x)}m_i^t + {}^{(d_x)}n_i^t\right)
\]

## 3. Multi-Head Extension

MHCADDI employs $K=8$ independent attention heads per layer. Each head has distinct projection matrices; their output vectors ($32$-dim each) are concatenated into a $256$-dim vector and then projected back to $32$ dimensions. This design enables modeling of multiple, potentially diverse, interaction types at the atom level, and enhances the expressiveness of the joint drug representation [1905.00534].

## 4. Readout and Drug-Pair Embedding

Following $T=3$ interleaved message-passing/co-attention layers, final atom representations are aggregated for each drug:
\[
d_x = \sum_{i \in d_x} f_r({}^{(d_x)}h_i^T)
\]
where $f_r$ is a single-layer LeakyReLU MLP ($32 \to 32$).

For prediction, drug pair embeddings are concatenated:
\[
[d_x \| d_y] \in \mathbb{R}^{64}
\]
serving as input for downstream scoring modules.

## 5. Prediction Heads and Training Procedure

### 5.1. Binary Classification (Per-Side Effect Ranking)

Input triplet: $(d_x, d_y, se_z)$, where $se_z$ is a one-hot vector over 964 side effect types. The matching score is defined as:
\[
f(d_x, d_y, se_z) = \|M_h\,d_x + se_z - M_t\,d_y\|_2^2 + \|M_h\,d_y + se_z - M_t\,d_x\|_2^2
\]
with $M_h, M_t \in \mathbb{R}^{32 \times 32}$.

Training minimizes the margin-based ranking loss:
\[
\mathcal{L} = \sum_{(d_x,d_y,se_z) \in \mathcal{P}} \sum_{(\tilde d_x, \tilde d_y) \notin \mathcal{P}} \max(0, \gamma - f(d_x,d_y,se_z) + f(\tilde d_x,\tilde d_y,se_z))
\]
where $\gamma=1$.

### 5.2. Multi-Label Classification (All Side Effects)

Predicts a 964-dimensional vector:
\[
y_{xy} = \sigma(W_p [d_x \| d_y] + b_p) \in (0,1)^{964}
\]
using per-label binary cross-entropy loss.

### 5.3. Hyperparameters and Optimization

- Layers ($T$): 3 interleaved message/co-attention
- Hidden dim: 32
- Attention heads: 8
- Dropout: 0.2 after each MLP or projection
- Optimizer: Adam, batch size 200, 30 epochs
- Learning rate: $\eta_t = 0.001 \times 0.96^{t\times10^{-6}}$
- Parameter initialization: Xavier uniform

MLP specifics:
- $f_i$: single-layer, no bias
- $f_v^t$: single-layer, no bias
- $f_e^t$: two-layer LeakyReLU, each $32\to32$
- $f_o^t$: single-layer LeakyReLU ($256 \to 32$)
- $f_r$: single-layer LeakyReLU ($32 \to 32$)
- $W_p$: $964 \times 64$, $b_p \in \mathbb{R}^{964}$

## 6. Relation to Recent Graph Neural Approaches

Recent architectures such as RGDA-DDI [2408.15310] use deeper residual GAT stacks, dual-attention fusion blocks, and multi-scale hierarchical pooling to increase representational power. In contrast, MHCADDI employs parallel multi-head co-attentive message-passing, integrating pairwise context at the atom level but does not implement: (a) separate substructure/global-structure GNN stacks, (b) hierarchical (layer-wise) SAGPooling, or (c) explicit dual (drug–drug and drug–DDP) attention for fusion across multiple feature spaces. In RGDA-DDI, these enhancements yielded improvements in metrics such as AUC and F1-score on large-scale DDI datasets, suggesting that while MHCADDI introduced the key paradigm of multi-head co-attentional fusion, further architectural depth and explicit dual-attention mechanisms can further improve predictive accuracy [2408.15310].

## 7. Practical Considerations and Implementation Notes

MHCADDI is suitable for end-to-end learning on large DDI datasets. All message-passing and attention operations are implemented using small feedforward networks with standard non-linearities (LeakyReLU), and attention heads scale linearly with $K$. Layer normalization and residual updates mitigate representation drift due to deep architectures. Drug graphs should be preprocessed from PubChem/SMILES with atomic and bond features as described. The architectural modularity of MHCADDI allows adaptation to different molecular feature types and side-effect ontologies by modifying input encodings or adjusting the number of output labels [1905.00534].

Source: https://www.emergentmind.com/topics/multi-head-co-attentive-drug-drug-interaction-encoder-mhcaddi