---
title: 'AudioRWKV: Efficient Audio Modeling'
url: https://www.emergentmind.com/topics/audiorwkv-a-rwkv
type: topic
---

# AudioRWKV: Efficient Audio Modeling

Searching arXiv for the specified AudioRWKV and related RWKV-ASR papers to ground the article in current paper metadata and citations.
{"query":"id:2509.02167 OR ti:\"AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition\"","max_results":5,"sort_by":"relevance"}
{"query":"id:2309.14758 OR ti:\"Exploring RWKV for Memory Efficient and Low Latency Streaming ASR\"","max_results":5,"sort_by":"relevance"}
AudioRWKV, abbreviated A-RWKV, is an audio modeling architecture built on the RWKV family and described as a highly efficient and stable architecture for audio modeling. It inherits the stable and efficient recurrent formulation of RWKV7, replaces the original 1D token-shift operation with a 2D depthwise separable convolution to better capture local spectro-temporal patterns, and adapts the original causal WKV kernel into a bidirectional WKV kernel (Bi-WKV) for global context modeling over the entire audio sequence while maintaining linear computational complexity [2509.02167]. Within the broader RWKV lineage, it occupies a different point in the design space from earlier streaming ASR applications of RWKV, which emphasized zero additional latency and single-frame state caching under strict online constraints [2309.14758].

## 1. Definition and lineage

A-RWKV is an adaptation of RWKV for audio spectrogram input, which is inherently 2D in time and frequency. The original RWKV model is presented as a hybrid that combines the sequential, linear-complexity nature of RNNs with the parallelizable block structure and representational power of Transformers. In its original NLP form, RWKV employs a 1D token shift operation to inject local context; A-RWKV modifies that mechanism for spectrogram modeling by using 2D local mixing [2509.02167].

The architectural lineage is important because RWKV had already been explored for audio in a different setting: streaming automatic speech recognition. In that earlier work, RWKV was proposed for streaming ASR because full-sequence attention is non-streamable and computationally expensive, whereas RWKV combines the superior performance of transformers and the inference efficiency of RNNs. The streaming ASR formulation used no future context, had zero additional latency, and required only a single frame’s state to be cached at inference [2309.14758].

This suggests a bifurcation within audio RWKV research. One branch uses causal recurrence for low-latency streaming ASR; the other, represented by A-RWKV, uses bidirectional recurrence for audio pattern recognition over full clips. The distinction is architectural rather than nominal: both inherit RWKV recurrence, but they optimize for different operating regimes.

## 2. Architectural composition

A-RWKV takes as input a mel spectrogram
$$
S \in \mathbb{R}^{B \times 1 \times H \times W}.
$$
An initial Conv2D layer projects the spectrogram into a sequence of patch features, followed by flattening and addition of a learnable positional embedding:
$$
x_0 = \mathrm{Flatten}(\mathrm{Conv2D}(S)) + \mathbf{P}_{pos}.
$$
The model then stacks \(N\) identical blocks consisting of bidirectional spatial mixing and channel mixing, with pre-LayerNorm and residuals, followed by global pooling and a classification layer [2509.02167].

The defining adaptation relative to standard RWKV is the replacement of the 1D token shift with a 2D Depthwise Separable Convolution. With the 2D input reshaped from the sequence, the local residual is
$$
x_{res} = \mathrm{Flatten}(\mathrm{DWConv2D}(X) - X).
$$
This local residual is used to condition dynamic parameters in the recurrent or global mixing stage. The details specify the conditioned form as
$$
x^{\square}_{t} = x_t + x_{res,t} \odot \mathbf{\mu}_{\square}.
$$
The purpose of this substitution is explicit: the 2D depthwise separable convolution better captures local spectro-temporal patterns inherent to audio [2509.02167].

The model therefore combines two inductive biases. The first is local spectro-temporal structure through DWConv2D. The second is long-range sequence aggregation through RWKV-style recurrent WKV updates. A plausible implication is that A-RWKV is designed to avoid the weak local inductive bias attributed in the comparison table to AST and Audio Mamba while retaining linear-complexity sequence processing.

## 3. Bidirectional WKV and global context

The core sequence operator in A-RWKV is a bidirectional extension of the RWKV7 WKV recurrence. The details summarize the recurrent update abstractly as
$$
\mathbf{wkv}_t = \mathbf{wkv}_{t-1} f(w_t, \kappa_t, a_t) + g(v_t, \tilde{k}_t),
$$
where \(w_t\) is a per-time decay, \(k_t\) and \(v_t\) are key and value terms, and all parameters are learned and can be dynamically modulated by the DWConv2D local residuals [2509.02167].

A-RWKV extends this recurrence bidirectionally by scanning both forward and backward along the sequence, producing \(p^{\rightarrow}\) and \(p^{\leftarrow}\). A learnable fusion gate \(\mathbf{G}\), derived from the local residual, combines the two directions:
$$
p_{fused} = \mathbf{G} \odot p^{\rightarrow} + (1 - \mathbf{G}) \odot \mathrm{Flip}(p^{\leftarrow}, \mathrm{dim}=1).
$$
This dynamically weights forward and backward information at each location [2509.02167].

A common misconception is that linear-complexity recurrent models necessarily forfeit full-sequence context. In A-RWKV, the bidirectional WKV kernel is introduced precisely to enable global context modeling over the entire audio sequence while maintaining linear computational complexity. The paper contrasts this with standard attention, which is quadratic in sequence length \(O(L^2)\), and presents A-RWKV as supporting arbitrarily long input with \(O(L)\) sequence complexity [2509.02167].

The classification readout is given as
$$
Y_{pred} = \mathrm{FC}(\mathrm{Mean}(\mathrm{LayerNorm}(x_N), \mathrm{dim}=1)).
$$
This positions A-RWKV as a full-sequence audio backbone rather than a purely autoregressive decoder.

## 4. Stability, scaling, and efficiency

The stability claim in A-RWKV rests on the RWKV7 foundation. The paper attributes to RWKV7 a highly stable dynamic state evolution with exponentially decaying memory and per-channel updates, and states that A-RWKV inherits this, experimentally enabling scaling to much larger model sizes and data without instability [2509.02167]. In the same comparison, Audio Mamba is described as becoming unstable at large scale, whereas AST is described as struggling with memory and computational scaling due to quadratic cost.

The efficiency claims are stated in both algorithmic and empirical terms. Algorithmically, A-RWKV has linear sequence complexity \(O(L)\), in contrast to AST’s \(O(L^2)\). Empirically, for long-form audio of approximately 5 minutes 28 seconds, WKV7 achieves up to a 13.3X speedup in processing, and A-RWKV’s kernel is reported as up to 13.3× faster than AST with FlashAttention at sequence length approximately 2048 [2509.02167]. The details further state that AST runs out of memory on medium-to-long sequences, whereas A-RWKV scales to 5-minute+ audio with stable memory and nearly flat throughput.

The model family is also presented as scalable in size. The detailed benchmark table lists A-RWKV-T at 6M parameters, A-RWKV-S at 23M, and A-RWKV-B at 91M. The abstract, however, states that A-RWKV-S is 22M and achieves performance parity with AuM-B at 92M while exhibiting more stable throughput than AST [2509.02167]. The coexistence of 22M in the abstract and 23M in the detailed table is part of the reported record.

| Model | Params | Selected from-scratch results |
|---|---:|---|
| AST-B | 86M | AudioSet2M 35.23 mAP; VGGSound 39.88; ESC-50 74.2 |
| AuM-B | 92M | AudioSet2M 32.43 mAP; VGGSound 42.58; ESC-50 — |
| A-RWKV-T | 6M | AudioSet2M 30.07 mAP; VGGSound 39.82; ESC-50 74.6 |
| A-RWKV-S | 23M | AudioSet2M 37.84 mAP; VGGSound 43.15; ESC-50 77.5 |
| A-RWKV-B | 91M | AudioSet2M 40.91 mAP; VGGSound 45.37; ESC-50 80.4 |

These results are accompanied by the summary claim that, from scratch, A-RWKV outperforms AST and AuM across all size regimes and datasets, and that smaller A-RWKV models can match or outperform significantly larger attention- or SSM-based models [2509.02167].

## 5. Empirical behavior across tasks

The evaluation suite spans AudioSet, VGGSound, ESC-50, SpeechCommands V2, and NSynth Pitch. The reported metrics are mean Average Precision for AudioSet and accuracy for the other tasks [2509.02167]. The model is therefore positioned as a general-purpose audio pattern recognition backbone rather than a task-specific speech recognizer.

The detailed from-scratch results emphasize scaling behavior. On AudioSet2M, A-RWKV-B reaches 40.91 mAP, compared with 35.23 for AST-B and 32.43 for AuM-B. On VGGSound, A-RWKV-B reaches 45.37 accuracy, compared with 39.88 for AST-B and 42.58 for AuM-B. On ESC-50, A-RWKV-B reaches 80.4, compared with 74.2 for AST-B [2509.02167].

The fine-tuning results after AudioSet2M pre-training follow the same pattern. The paper reports VGGSound at 48.91% for A-RWKV-B, compared with 46.61% for AuM-B and 44.17% for AST-B; ESC-50 at 86.8% for A-RWKV-B, compared with 83.5% for AST-B; and SpeechCommands V2 at 96.83% for A-RWKV-B, compared with 94.78% for AuM-B and 90.37% for AST-B [2509.02167].

The ablation study on AudioSet2M mAP isolates the contribution of the major components. Starting from RWKV7 with causal 1D shift at 34.50, adding Bi-Scan yields 38.39, adding Fusion Gate yields 39.02, adding Q-Shift (2D shift) yields 39.35, and the full A-RWKV with ConvShift+Gate reaches 40.91 [2509.02167]. This supports the paper’s explicit takeaway that the combination of bidirectional WKV, fusion gating, and 2D convolutional local mixing is crucial to performance.

## 6. Relation to streaming RWKV for ASR, scope, and limitations

A-RWKV should not be conflated with RWKV-Transducer models developed for streaming ASR. In the streaming ASR work, the input was 80-dim filterbank features processed with a convolutional subsampling layer followed by RWKV blocks; the key streaming properties were no future context, latency of 0, left context of 1, and minimal inference memory because only the last state needed to be cached [2309.14758]. That work showed that RWKV-Transducer and RWKV-Boundary-Aware-Transducer achieved accuracy comparable to or better than chunk conformer transducer under severe latency and memory constraints.

By contrast, A-RWKV is explicitly bidirectional and is evaluated on audio pattern recognition benchmarks rather than transducer-based ASR. The two lines of work therefore address different objectives: one prioritizes online recognition with zero additional latency, and the other prioritizes efficient full-sequence bidirectional modeling with global context [2509.02167]. This suggests that “audio RWKV” is best understood as a family of adaptations rather than a single fixed architecture.

The strengths and limitations reported for A-RWKV are correspondingly specific. Its stated strengths are linear computational complexity for end-to-end sequence modeling, stable training and scalability, parameter efficiency, and a balance between local and global modeling through 2D convolutional shift and Bi-WKV [2509.02167]. Its use cases are environmental sounds, speech commands, music, and event detection, especially long-form audio scenarios where global context and local details matter and efficiency is critical. The principal comparative limitations discussed in the source material are external rather than internal: AST is described as impractical on long-form audio because of quadratic scaling and out-of-memory behavior, while AuM is described as fragile or unstable at scale. A-RWKV is positioned as an alternative to both, rather than as a universal replacement for causal streaming encoders.

Source: https://www.emergentmind.com/topics/audiorwkv-a-rwkv