Papers
Topics
Authors
Recent
Search
2000 character limit reached

DyKen-Hyena: Dynamic Kernel Generation for MIR

Updated 10 July 2026
  • DyKen-Hyena is a multimodal intent recognition architecture that reframes fusion by modulating text processing using dynamic, per-token convolutions.
  • It generates per-token convolutional kernels from aligned audio-visual cues via cross-modal attention to condition early textual feature extraction.
  • The integration of short dynamic convolutions with FFT-based long convolution leads to robust performance across intent recognition benchmarks.

DyKen-Hyena is a multimodal intent recognition architecture that reframes multimodal fusion as processing modulation rather than representation-level feature fusion. Introduced in "DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition" (Wang et al., 12 Sep 2025), it uses cross-modal attention to translate aligned audio-visual context into dynamic, per-token convolutional kernels that modulate textual feature extraction inside a Hyena-based module placed between the initial word embedding layer and a BERT encoder. The model is designed for settings in which language carries explicit content while prosody and facial expression provide intent-relevant but temporally localized cues; its central claim is that non-verbal signals should modulate how text is processed rather than be directly added back to text representations.

1. Problem formulation and conceptual shift

Multimodal Intent Recognition (MIR) aims to infer speaker intent from synchronized language, audio, and video streams. In this setting, language provides the explicit content, but the true intent is often carried by non-verbal signals such as intonation, prosody, and facial expression. A central difficulty is that audio-visual streams can be intent-irrelevant or even conflicting at different moments, so naïve fusion can corrupt the primary linguistic features (Wang et al., 12 Sep 2025).

DyKen-Hyena is explicitly positioned against prior approaches that combine strong unimodal encoders through cross-modal attention, gating, tensor fusion, or injection into a unified Transformer after unimodal abstraction. The paper argues that such late or intermediate fusion “risk[s] corrupting the primary linguistic features with noisy or irrelevant non-verbal signals,” because it does not model how non-verbal cues color specific words or short phrases in real time. DyKen-Hyena therefore shifts MIR from feature fusion to processing modulation: audio-visual signals are treated as directives that condition the operator applied to text at an early stage of representation learning.

This distinction is not merely terminological. In DyKen-Hyena, the multimodal branch does not directly overwrite or augment textual features; instead, it parameterizes token-local convolutions that shape how textual information is extracted. The paper presents this as a mechanism for preserving linguistic integrity while injecting non-verbal nuance at token granularity.

2. Architectural composition

The model comprises a BERT-based textual backbone, aligned audio and visual feature streams, token-level cross-modal attention blocks, a lightweight kernel generator, and a Hyena operator adapted to combine dynamic short convolutions with long-range FFT-based convolution (Wang et al., 12 Sep 2025).

Let text features be Ft∈RLt×DtF_t \in \mathbb{R}^{L_t \times D_t}, audio features Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}, and visual features Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}. Audio and visual sequences are temporally aligned to text tokens, yielding aligned sequences of length LtL_t. These aligned non-text features are projected to DtD_t and concatenated into

Fav∈RLt×2Dt.F_{av} \in \mathbb{R}^{L_t \times 2D_t}.

Text tokens then act as queries over aligned audio-visual features as keys and values:

Cattn=MultiHeadAttn(Q=Ft,K=Fav,V=Fav),C_{\text{attn}} = \mathrm{MultiHeadAttn}(Q = F_t, K = F_{av}, V = F_{av}),

with the standard single-head form

Attn(Q,K,V)=Softmax(QKT/d)V.\mathrm{Attn}(Q, K, V) = \mathrm{Softmax}(QK^T / \sqrt{d})V.

A residual connection and layer normalization produce the local multimodal context:

Clocal=LayerNorm(Cattn+Ft).C_{\text{local}} = \mathrm{LayerNorm}(C_{\text{attn}} + F_t).

This module is inserted between the initial word embedding layer and the main BERT encoder. The resulting fused embeddings propagate through the BERT stack for higher-level contextualization. The placement is significant because the paper treats DyKen-Hyena as a pre-encoder fusion module rather than a late-stage fusion head.

Temporal alignment is structurally essential. The model assumes that for each token position t∈{1,…,Lt}t \in \{1,\dots,L_t\}, aligned audio-visual features can be obtained so that kernel generation reflects the non-verbal context of that specific word or short phrase. The paper identifies this alignment as necessary for meaningful per-token modulation.

3. Dynamic kernel generation and Hyena modulation

The kernel generator is a small MLP that maps token-level multimodal context into per-token, per-channel short convolutional filters:

Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}0

These parameters are reshaped into

Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}1

where the short kernel size Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}2 is chosen from Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}3 and Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}4 is the primary configuration (Wang et al., 12 Sep 2025).

Let Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}5 denote the text features on the short-convolution path. For each token Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}6 and channel Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}7, the dynamically generated kernel Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}8 is applied over a local window centered at Fa∈RLa×DaF_a \in \mathbb{R}^{L_a \times D_a}9. With half-window Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}0, the operation is

Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}1

In implementation terms, an unfold operation constructs local windows, after which element-wise products with per-token kernels are summed.

DyKen-Hyena integrates this dynamic short convolution into a Hyena sequence module with three components. First, the locally modulated representation Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}2 is obtained by applying Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}3 to the Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}4 path. Second, a parallel path Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}5 gates Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}6 element-wise. Third, the gated representation is passed through a long convolution implemented with FFT, which provides a global receptive field with sub-quadratic complexity. The final fused representation is reintegrated through

Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}7

The paper’s distinction from both attention and static convolution is precise. Unlike standard attention, DyKen-Hyena uses attention only to derive token-level context vectors; these vectors do not directly augment text features, but parameterize the operator applied to them. Unlike static convolutions, the short filters are generated dynamically for each token and each channel, conditioned on aligned audio-visual cues. The result is a hybrid operator in which local modulation is dynamic and global modulation remains an FFT-based long convolution.

4. Training objective, benchmarks, and evaluation protocol

DyKen-Hyena is trained for intent classification with a standard softmax cross-entropy loss. For a mini-batch of Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}8 examples with labels Fv∈RLv×DvF_v \in \mathbb{R}^{L_v \times D_v}9 and predicted class probabilities LtL_t0, the loss is

LtL_t1

The paper does not introduce an additional out-of-scope-specific training loss or calibration strategy; out-of-scope robustness is described as arising from the learned representation itself (Wang et al., 12 Sep 2025).

The evaluation uses two MIR benchmarks. MIntRec is a closed-set classification benchmark. MIntRec2.0 includes open intent detection and reports OID Accuracy, F1-IS, F1-OOS, and OID F1. Across experiments, reported numbers are averaged over 5 random seeds. The backbone is bert-base-uncased, and the audio-visual streams are aligned to text token length LtL_t2 using standard alignment techniques before projection and concatenation into LtL_t3.

The paper reports the following main benchmark outcomes.

Benchmark DyKen-Hyena results Comparative statement
MIntRec2.0 WF=59.79, WP=60.21, F1=55.02, Prec=58.35, Rec=54.53, oid_acc=45.63, F1-IS=46.21, F1-OOS=38.69, oid_f1=45.97 +10.46% absolute improvement in F1-OOS over MAG-BERT
MIntRec Acc=73.66, F1=69.26, Prec=70.30, Rec=69.54, WF=73.05, WP=73.30 New SOTA with +1.84% Acc over best baselines

The evaluation setup also clarifies what is not reported. The paper does not provide exact parameter counts, throughput, or latency, and does not include a repository link or license information in the supplied content. For exact replication, the alignment procedure, audio-visual feature extraction details, and training hyperparameters such as learning rate and batch size are stated to be absent from the text.

5. Empirical behavior, ablations, and robustness

The principal empirical claim is that DyKen-Hyena achieves state-of-the-art results on both MIntRec and MIntRec2.0, with especially pronounced gains in out-of-scope detection. On MIntRec2.0, the model improves F1-OOS by +10.46% absolute over MAG-BERT, from 28.23 to 38.69. On MIntRec, it reports Acc=73.66, F1=69.26, Prec=70.30, Rec=69.54, WF=73.05, and WP=73.30, with improvements over the best baselines of +1.84%, +0.79%, +0.57%, +0.98%, +1.47%, and +1.15%, respectively (Wang et al., 12 Sep 2025).

The ablation study identifies the dynamic short convolution as the core contributor to robustness. Removing it (w/o-DynamicShortConv) yields the largest degradation: F1-OOS falls from 38.69% to 22.31%, and overall F1 falls from 55.02% to 52.82%. Removing cross-modal attention (w/o-Attention) lowers F1-OOS to 34.51% and F1 to 54.31%. Removing the long convolution (w/o-LongConv) produces F1-OOS=31.93% and F1=53.88%. The pattern supports the paper’s interpretation that dynamic local kernels and static long-range modeling play complementary roles.

Kernel size analysis further sharpens this interpretation. On MIntRec2.0, LtL_t4 yields F1-OOS=38.69% and outperforms LtL_t5 on open-set detection, where F1-OOS is 28.76%. On MIntRec, LtL_t6 is reported as best across the board. The paper interprets this as evidence that phrasal-level modulation—token plus neighbors—better matches how non-verbal cues color short phrases, whereas larger kernels such as LtL_t7 may introduce irrelevant context.

Qualitatively, the model shows notable gains on intents that rely heavily on non-verbal signals, including Emphasize, Joke, and Taunt. The paper attributes this to the model’s ability to associate vocal and facial cues with otherwise ambiguous text. More generally, its robustness argument has three parts: it prevents direct corruption of textual features, it applies token-specific modulation aligned with local non-verbal context, and it combines local dynamic kernels with a static global long convolution to construct a tighter in-scope manifold.

6. Computational profile, limitations, and relation to broader Hyena research

DyKen-Hyena uses Hyena’s long convolution via FFT to obtain near-linear runtime with a global receptive field, while adding moderate overhead from cross-modal attention and dynamic short convolution. The paper states that it remains more efficient than full attention-based fusion across many layers, but exact parameter counts and throughput or latency are not reported (Wang et al., 12 Sep 2025).

Its limitations are correspondingly concrete. The model incurs a moderate increase in computational cost from the attention and dynamic convolution steps. It relies on explicit feature alignment, and the paper points to end-to-end handling of unaligned streams as a future direction. The kernel generator is a simple MLP, and the authors note that more expressive generators or alternative conditional operators, including richer dynamic filters or state-space formulations, may further improve performance.

Within the broader Hyena literature, DyKen-Hyena is specific to multimodal intent recognition and should be distinguished from other Hyena-based lines of work. Other papers use Hyena for long convolution sequence modeling and inference acceleration (Oncescu et al., 2024), PDE operators (Patil et al., 2023), or transformer distillation into long convolution models (Ralambomihanta et al., 2024), but do not define DyKen-Hyena. This broader context matters because DyKen-Hyena inherits the Hyena idea of combining multiplicative or gated local processing with long-range convolution, while specializing it to multimodal token-level modulation.

A plausible systems implication is that DyKen-Hyena’s FFT-based long convolution could benefit from later Hyena inference optimizations if deployed in long-sequence regimes. "Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond" states that for Hyena-class long convolution sequence models, exact inference can be reduced to quasilinear LtL_t8 time, and that with data-dependent kernels an exact extension remains possible at the same asymptotic complexity but with about LtL_t9 extra FLOPs relative to the data-independent case (Oncescu et al., 2024). This does not constitute a reported DyKen-Hyena result, but it suggests that the model’s static long-convolution component is compatible with a maturing systems stack for Hyena-family operators.

Taken together, DyKen-Hyena is best understood as a Hyena-based MIR architecture in which cross-modal attention is used not for direct fusion, but for conditional operator generation. Its defining property is that non-verbal context is converted into per-token convolutional kernels that modulate early textual processing, while an FFT-based long convolution preserves global dependency modeling. The empirical record supplied in the paper supports this design chiefly through improved out-of-scope detection and through ablations showing that the dynamic short convolution is the dominant source of the model’s robustness.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DyKen-Hyena.