---
title: 'DyKen-Hyena: Dynamic Kernel Generation for MIR'
url: https://www.emergentmind.com/topics/dyken-hyena
type: topic
---

# DyKen-Hyena: Dynamic Kernel Generation for MIR

DyKen-Hyena is a multimodal intent recognition architecture that reframes multimodal fusion as **processing modulation** rather than representation-level feature fusion. Introduced in "DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition" [2509.09940], it uses cross-modal attention to translate aligned audio-visual context into **dynamic, per-token convolutional kernels** that modulate textual feature extraction inside a Hyena-based module placed between the initial word embedding layer and a BERT encoder. The model is designed for settings in which language carries explicit content while prosody and facial expression provide intent-relevant but temporally localized cues; its central claim is that non-verbal signals should modulate how text is processed rather than be directly added back to text representations.

## 1. Problem formulation and conceptual shift

Multimodal Intent Recognition (MIR) aims to infer speaker intent from synchronized language, audio, and video streams. In this setting, language provides the explicit content, but the true intent is often carried by non-verbal signals such as intonation, prosody, and facial expression. A central difficulty is that audio-visual streams can be intent-irrelevant or even conflicting at different moments, so naïve fusion can corrupt the primary linguistic features [2509.09940].

DyKen-Hyena is explicitly positioned against prior approaches that combine strong unimodal encoders through cross-modal attention, gating, tensor fusion, or injection into a unified Transformer after unimodal abstraction. The paper argues that such late or intermediate fusion “risk[s] corrupting the primary linguistic features with noisy or irrelevant non-verbal signals,” because it does not model how non-verbal cues color specific words or short phrases in real time. DyKen-Hyena therefore shifts MIR from **feature fusion** to **processing modulation**: audio-visual signals are treated as directives that condition the operator applied to text at an early stage of representation learning.

This distinction is not merely terminological. In DyKen-Hyena, the multimodal branch does not directly overwrite or augment textual features; instead, it parameterizes token-local convolutions that shape how textual information is extracted. The paper presents this as a mechanism for preserving linguistic integrity while injecting non-verbal nuance at token granularity.

## 2. Architectural composition

The model comprises a BERT-based textual backbone, aligned audio and visual feature streams, token-level cross-modal attention blocks, a lightweight kernel generator, and a Hyena operator adapted to combine dynamic short convolutions with long-range FFT-based convolution [2509.09940].

Let text features be $F_t \in \mathbb{R}^{L_t \times D_t}$, audio features $F_a \in \mathbb{R}^{L_a \times D_a}$, and visual features $F_v \in \mathbb{R}^{L_v \times D_v}$. Audio and visual sequences are temporally aligned to text tokens, yielding aligned sequences of length $L_t$. These aligned non-text features are projected to $D_t$ and concatenated into
$$
F_{av} \in \mathbb{R}^{L_t \times 2D_t}.
$$
Text tokens then act as queries over aligned audio-visual features as keys and values:
$$
C_{\text{attn}} = \mathrm{MultiHeadAttn}(Q = F_t, K = F_{av}, V = F_{av}),
$$
with the standard single-head form
$$
\mathrm{Attn}(Q, K, V) = \mathrm{Softmax}(QK^T / \sqrt{d})V.
$$
A residual connection and layer normalization produce the local multimodal context:
$$
C_{\text{local}} = \mathrm{LayerNorm}(C_{\text{attn}} + F_t).
$$

This module is inserted between the initial word embedding layer and the main BERT encoder. The resulting fused embeddings propagate through the BERT stack for higher-level contextualization. The placement is significant because the paper treats DyKen-Hyena as a **pre-encoder fusion module** rather than a late-stage fusion head.

Temporal alignment is structurally essential. The model assumes that for each token position $t \in \{1,\dots,L_t\}$, aligned audio-visual features can be obtained so that kernel generation reflects the non-verbal context of that specific word or short phrase. The paper identifies this alignment as necessary for meaningful per-token modulation.

## 3. Dynamic kernel generation and Hyena modulation

The kernel generator is a small MLP that maps token-level multimodal context into per-token, per-channel short convolutional filters:
$$
K_{\text{params}} = \mathrm{MLP}(C_{\text{local}}) \in \mathbb{R}^{L_t \times (D_t \cdot K_s)}.
$$
These parameters are reshaped into
$$
F_{\text{short}} \in \mathbb{R}^{L_t \times D_t \times K_s},
$$
where the short kernel size $K_s$ is chosen from $\{1,3,5\}$ and $K_s=3$ is the primary configuration [2509.09940].

Let $X \in \mathbb{R}^{L_t \times D_t}$ denote the text features on the short-convolution path. For each token $t$ and channel $d$, the dynamically generated kernel $k_t^d \in \mathbb{R}^{K_s}$ is applied over a local window centered at $t$. With half-window $m = \lfloor K_s/2 \rfloor$, the operation is
$$
y_t^d = \sum_{i=-m}^{m} k_t^d[i] \cdot x_{t+i}^d.
$$
In implementation terms, an unfold operation constructs local windows, after which element-wise products with per-token kernels are summed.

DyKen-Hyena integrates this dynamic short convolution into a Hyena sequence module with three components. First, the locally modulated representation $X_{\text{conv}}$ is obtained by applying $F_{\text{short}}$ to the $X_1$ path. Second, a parallel path $X_2$ gates $X_{\text{conv}}$ element-wise. Third, the gated representation is passed through a long convolution implemented with FFT, which provides a global receptive field with sub-quadratic complexity. The final fused representation is reintegrated through
$$
F_{\text{final}} = \mathrm{LayerNorm}(F_{\text{fusion}} + F_t).
$$

The paper’s distinction from both attention and static convolution is precise. Unlike standard attention, DyKen-Hyena uses attention only to derive token-level context vectors; these vectors do not directly augment text features, but parameterize the operator applied to them. Unlike static convolutions, the short filters are generated dynamically for each token and each channel, conditioned on aligned audio-visual cues. The result is a hybrid operator in which **local modulation is dynamic** and **global modulation remains an FFT-based long convolution**.

## 4. Training objective, benchmarks, and evaluation protocol

DyKen-Hyena is trained for intent classification with a standard softmax cross-entropy loss. For a mini-batch of $N$ examples with labels $y_i$ and predicted class probabilities $p(y \mid x_i)$, the loss is
$$
L_{ce} = - \frac{1}{N} \sum_{i=1}^{N} \log p(y_i \mid x_i).
$$
The paper does not introduce an additional out-of-scope-specific training loss or calibration strategy; out-of-scope robustness is described as arising from the learned representation itself [2509.09940].

The evaluation uses two MIR benchmarks. **MIntRec** is a closed-set classification benchmark. **MIntRec2.0** includes open intent detection and reports OID Accuracy, F1-IS, F1-OOS, and OID F1. Across experiments, reported numbers are averaged over 5 random seeds. The backbone is `bert-base-uncased`, and the audio-visual streams are aligned to text token length $L_t$ using standard alignment techniques before projection and concatenation into $F_{av}$.

The paper reports the following main benchmark outcomes.

| Benchmark | DyKen-Hyena results | Comparative statement |
|---|---:|---|
| MIntRec2.0 | WF=59.79, WP=60.21, F1=55.02, Prec=58.35, Rec=54.53, oid_acc=45.63, F1-IS=46.21, F1-OOS=38.69, oid_f1=45.97 | +10.46% absolute improvement in F1-OOS over MAG-BERT |
| MIntRec | Acc=73.66, F1=69.26, Prec=70.30, Rec=69.54, WF=73.05, WP=73.30 | New SOTA with +1.84% Acc over best baselines |

The evaluation setup also clarifies what is not reported. The paper does not provide exact parameter counts, throughput, or latency, and does not include a repository link or license information in the supplied content. For exact replication, the alignment procedure, audio-visual feature extraction details, and training hyperparameters such as learning rate and batch size are stated to be absent from the text.

## 5. Empirical behavior, ablations, and robustness

The principal empirical claim is that DyKen-Hyena achieves state-of-the-art results on both MIntRec and MIntRec2.0, with especially pronounced gains in out-of-scope detection. On MIntRec2.0, the model improves F1-OOS by +10.46% absolute over MAG-BERT, from 28.23 to 38.69. On MIntRec, it reports Acc=73.66, F1=69.26, Prec=70.30, Rec=69.54, WF=73.05, and WP=73.30, with improvements over the best baselines of +1.84%, +0.79%, +0.57%, +0.98%, +1.47%, and +1.15%, respectively [2509.09940].

The ablation study identifies the dynamic short convolution as the core contributor to robustness. Removing it (`w/o-DynamicShortConv`) yields the largest degradation: F1-OOS falls from 38.69% to 22.31%, and overall F1 falls from 55.02% to 52.82%. Removing cross-modal attention (`w/o-Attention`) lowers F1-OOS to 34.51% and F1 to 54.31%. Removing the long convolution (`w/o-LongConv`) produces F1-OOS=31.93% and F1=53.88%. The pattern supports the paper’s interpretation that dynamic local kernels and static long-range modeling play complementary roles.

Kernel size analysis further sharpens this interpretation. On MIntRec2.0, $K_s=3$ yields F1-OOS=38.69% and outperforms $K_s=1$ on open-set detection, where F1-OOS is 28.76%. On MIntRec, $K_s=3$ is reported as best across the board. The paper interprets this as evidence that phrasal-level modulation—token plus neighbors—better matches how non-verbal cues color short phrases, whereas larger kernels such as $K_s=5$ may introduce irrelevant context.

Qualitatively, the model shows notable gains on intents that rely heavily on non-verbal signals, including **Emphasize**, **Joke**, and **Taunt**. The paper attributes this to the model’s ability to associate vocal and facial cues with otherwise ambiguous text. More generally, its robustness argument has three parts: it prevents direct corruption of textual features, it applies token-specific modulation aligned with local non-verbal context, and it combines local dynamic kernels with a static global long convolution to construct a tighter in-scope manifold.

## 6. Computational profile, limitations, and relation to broader Hyena research

DyKen-Hyena uses Hyena’s long convolution via FFT to obtain near-linear runtime with a global receptive field, while adding moderate overhead from cross-modal attention and dynamic short convolution. The paper states that it remains more efficient than full attention-based fusion across many layers, but exact parameter counts and throughput or latency are not reported [2509.09940].

Its limitations are correspondingly concrete. The model incurs a moderate increase in computational cost from the attention and dynamic convolution steps. It relies on explicit feature alignment, and the paper points to end-to-end handling of unaligned streams as a future direction. The kernel generator is a simple MLP, and the authors note that more expressive generators or alternative conditional operators, including richer dynamic filters or state-space formulations, may further improve performance.

Within the broader Hyena literature, DyKen-Hyena is specific to multimodal intent recognition and should be distinguished from other Hyena-based lines of work. Other papers use Hyena for long convolution sequence modeling and inference acceleration [2410.12982], PDE operators [2306.16524], or transformer distillation into long convolution models [2401.17574], but do not define DyKen-Hyena. This broader context matters because DyKen-Hyena inherits the Hyena idea of combining multiplicative or gated local processing with long-range convolution, while specializing it to multimodal token-level modulation.

A plausible systems implication is that DyKen-Hyena’s FFT-based long convolution could benefit from later Hyena inference optimizations if deployed in long-sequence regimes. "Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond" states that for Hyena-class long convolution sequence models, exact inference can be reduced to quasilinear $O(L\log^2L)$ time, and that with data-dependent kernels an exact extension remains possible at the same asymptotic complexity but with about $2\times$ extra FLOPs relative to the data-independent case [2410.12982]. This does not constitute a reported DyKen-Hyena result, but it suggests that the model’s static long-convolution component is compatible with a maturing systems stack for Hyena-family operators.

Taken together, DyKen-Hyena is best understood as a Hyena-based MIR architecture in which cross-modal attention is used not for direct fusion, but for conditional operator generation. Its defining property is that non-verbal context is converted into per-token convolutional kernels that modulate early textual processing, while an FFT-based long convolution preserves global dependency modeling. The empirical record supplied in the paper supports this design chiefly through improved out-of-scope detection and through ablations showing that the dynamic short convolution is the dominant source of the model’s robustness.

Source: https://www.emergentmind.com/topics/dyken-hyena