---
title: Bidirectional WKV Kernel (Bi-WKV)
url: https://www.emergentmind.com/topics/bidirectional-wkv-kernel-bi-wkv
type: topic
---

# Bidirectional WKV Kernel (Bi-WKV)

Bidirectional WKV Kernel (Bi-WKV) denotes a family of RWKV-derived sequence-mixing operators that extend the original causal WKV kernel so that each position incorporates information from both earlier and later positions. In recent work, the term is used in at least three closely related but non-identical ways: as a symmetric distance-decayed global aggregation in SATEM denoising, as the combination of forward and backward WKV scans in image restoration, and as a gated fusion of two causal RWKV7 streams in audio pattern recognition [2503.22223][2412.03814][2509.02167]. The common objective is to recover bidirectional context without reverting to quadratic self-attention.

## 1. Origins in the RWKV WKV kernel

Bi-WKV is derived from the original unidirectional WKV mechanism in RWKV. In the DREMnet formulation, a sequence \(x\in\mathbb R^{T\times C}\) is projected, after token shift, into \(R_t,K_t,V_t\in\mathbb R^C\), and the unidirectional output at position \(t\) is a normalized exponentially decayed sum over the prefix together with a learned current-token reward \(u\) [2503.22223]:
\[
\mathrm{WKV}_{t}
=
\frac{
\sum_{i=0}^{t-1}
\exp\bigl(k_i + (i-t+1)\,w\bigr)\,v_i
+
\exp\bigl(u + k_t\bigr)\,v_t
}{
\sum_{i=0}^{t-1}
\exp\bigl(k_i + (i-t+1)\,w\bigr)
+
\exp\bigl(u + k_t\bigr)
}.
\]
The same paper gives an \(O(T\!\cdot\!C)\) implementation through running numerators and denominators,
\[
N_t = e^{u + k_t}\,v_t + e^{-\,w}\circ N_{t-1},\qquad
D_t = e^{u + k_t} + e^{-\,w}\circ D_{t-1},\qquad
\mathrm{WKV}_t = \frac{N_t}{D_t},
\]
so the operator acts as a normalized exponentially weighted accumulator.

RWKV-IR describes the same underlying idea from the perspective of replacing \(O(T^2)\) dot-product attention by two \(O(T\,C)\) recurrences with exponential decay. There the forward scan maintains
\[
m_t^{+}=\alpha\odot m_{t-1}^{+}+ek_t\odot v_t,\qquad
n_t^{+}=\alpha\odot n_{t-1}^{+}+ek_t,\qquad
wkv_t^{+}=m_t^{+}/n_t^{+},
\]
with an analogous backward scan, followed by gating with receptance \(R\) [2412.03814]. AudioRWKV presents a RWKV7 version in which the recurrent state consists of two channel-wise running sums \(S_t\) and \(U_t\),
\[
S_t=w_t\odot S_{t-1}+k_t,\qquad
U_t=w_t\odot U_{t-1}+k_t\odot v_t,\qquad
\mathrm{WKV}_{\mathrm{causal}(x)_t}=r_t\odot(U_t\oslash S_t),
\]
where \(w_t\in(0,1)^D\) is a per-channel decay or forget gate [2509.02167].

These formulations differ in parameterization, but they share the same structural principle: WKV computes an attention-like weighted value aggregation through recurrently updated state rather than a full pairwise attention matrix. This suggests that Bi-WKV is best understood as a bidirectionalization of a recurrent normalized weighting operator rather than as a single canonical kernel.

## 2. Principal bidirectional formulations

The literature does not define a unique Bi-WKV equation. Instead, bidirectionality is introduced by different constructions that all permit future-to-past information flow.

| Work | Bidirectional construction | Merge or control |
|---|---|---|
| DREMnet | Full sum over all \(i\neq t\) with symmetric distance decay \(|t-i|\) | learned \(w,u\) |
| RWKV-IR | Forward and backward WKV scans | merge, e.g. average |
| AudioRWKV | Causal WKV on original and reversed sequence | learned gate \(g_t\) |

In DREMnet, unidirectional decay is replaced by a symmetric relative-distance decay
\[
\delta_{t,i}=-(|t-i|-1)\,w,
\]
yielding
\[
\mathrm{BiWKV}_t
=
\frac{
\sum_{\substack{i=0\\ i\neq t}}^{T-1}
\exp(\delta_{t,i}+k_i)\,v_i
+
\exp(u+k_t)\,v_t
}{
\sum_{\substack{i=0\\ i\neq t}}^{T-1}
\exp(\delta_{t,i}+k_i)
+
\exp(u+k_t)
}.
\]
Here \(w\) is a learned decay vector and \(u\) is a learned self-reward; the defining change is the substitution of \(|t-i|\) for the one-sided temporal offset, so each position receives context from both directions [2503.22223].

RWKV-IR defines “Bi-WKV” as exactly the combination of forward and backward WKV scans. The forward and backward recurrent outputs \(wkv_t^{+}\) and \(wkv_t^{-}\) are merged, for example by
\[
wkv_t=\frac{wkv_t^{+}+wkv_t^{-}}{2},
\]
and the result is then gated through \(\sigma(R_t)\odot wkv_t\) [2412.03814]. AudioRWKV uses the same directional duplication idea on RWKV7, but replaces fixed averaging by a learned elementwise gate,
\[
\mathrm{WKV}_{\mathrm{bi}(x)_t}
=
g_t\odot p_t^{\rightarrow}
+
(1-g_t)\odot p_t^{\leftarrow},\qquad
g_t=\sigma(W_gx_t+b_g),
\]
where \(p_t^{\rightarrow}\) is the causal WKV on the original sequence and \(p_t^{\leftarrow}\) is the corresponding output on the reversed sequence mapped back to the original indexing [2509.02167].

A common misconception is that “Bi-WKV” names a single standardized kernel. The published formulations instead show a design space whose members share bidirectional context aggregation and RWKV-style linear recurrences, but differ in whether bidirectionality is encoded by an explicit symmetric kernel, by scan averaging, or by gated directional fusion.

## 3. Coupling with local receptive-field mechanisms

Across the cited systems, Bi-WKV is not used in isolation. Each architecture couples global bidirectional aggregation with an explicitly local mechanism intended to preserve short-range structure.

In DREMnet, locality is supplied by Covering Embedding. For a raw one-dimensional signal \(y\in\mathbb R^T\), each token is formed from an overlapping window,
\[
t_t=[\,y_t; y_{t+1}; \ldots; y_{t+n-1}\,]\in\mathbb R^C,
\]
with zero-padding at the end, and stacking all such tokens yields \(X\in\mathbb R^{T\times C}\). The stated purpose is to retain strong local receptive fields akin to CNNs, so that each token contains its own sample together with its next \(n-1\) neighbors [2503.22223].

RWKV-IR similarly replaces Vision-RWKV’s original Q-Shift with a depth-wise convolution shift, DC-Shift:
\[
\mathrm{DC\!-\!Shift}(X)
=\mathrm{Conv}_{1\times1}\Bigl(
\mathrm{GeLU}\bigl(
\mathrm{DWConv}_{k\times k}(
\mathrm{GeLU}(\mathrm{Conv}_{1\times1}(X))
)
\bigr)
\Bigr).
\]
The reported motivation is to capture all \(k\times k\) neighbors rather than the restricted directional substitution used by Q-Shift [2412.03814]. AudioRWKV adopts an analogous strategy for spectrogram inputs by replacing the original 1D token-shift with a 2D depthwise-separable convolution, denoted ConvShift, applied on Mel-spectrogram patches; a local residual \(x_{\mathrm{res},t}\) is then injected into the parameter-generating pathways through
\[
x_t^{\square}=x_t+x_{\mathrm{res},t}\odot\mu_{\square}.
\]
The stated effect is to combine a local spectro-temporal bias with global Bi-WKV attention [2509.02167].

A plausible implication is that published Bi-WKV systems treat bidirectional recurrent aggregation as a global mixer whose performance depends materially on an auxiliary local inductive bias. The local operator differs by modality—overlapping windows for one-dimensional SATEM signals, depth-wise convolutional shift for images, and depthwise-separable spectro-temporal convolution for audio—but the architectural pattern is consistent.

## 4. Role inside full network architectures

In DREMnet, Bi-WKV is embedded in an interpretable decoupled representation learning framework for semi-airborne transient electromagnetic signal denoising. The encoder disentangles the input signal \(x\) into
\[
(Z_s,Z_n)=E(x),
\]
where \(Z_s\) are “content” factors corresponding to the underlying clean EM response and \(Z_n\) are “context” factors corresponding to task-irrelevant or nuisance information. Disentanglement is enforced by minimizing a mutual-information upper bound (CLUB) between \(Z_s\) and \(Z_n\), and for a clean reference \(s\), forcing \(Z_n(s)\approx 0\) via a KL–Gaussian prior. After Covering Embedding on both factors, the streams are fused and passed through DR-blocks in which Bi-WKV performs signal mixing and a channel-mixing sublayer performs per-token MLP-style transformation [2503.22223].

RWKV-IR places Bi-WKV inside a transformer-like restoration layer. The layer applies LayerNorm, DC-Shift, linear projections to \(R\), \(K\), and \(V\), then either a single Bi-WKV scan or the Cross-Bi-WKV extension, in which one Bi-WKV pass operates on row-major flattening and another on column-major flattening. Their outputs, \(wkv^h\) and \(wkv^v\), are fused as
\[
wkv_t^{\rm cross}=\frac12\,wkv_t^h+\frac12\,wkv_t^v.
\]
The output is then gated by \(R\), projected by \(W_O\), and combined with residual and ChannelMix sublayers [2412.03814].

AudioRWKV uses Bi-WKV in a sequence model built on RWKV7 for audio pattern recognition. Spectrograms are patch-embedded into a sequence, locally transformed by ConvShift, and then processed by bidirectional RWKV recurrences whose forward and backward outputs are fused by the learned gate \(g_t\). The paper explicitly positions this as a way to maintain the stable and efficient recurrent formulation of RWKV7 while adapting the original causal WKV kernel to global bidirectional context over the full audio sequence [2509.02167].

These uses show that Bi-WKV functions less as an isolated primitive than as the sequence-mixing core of broader hybrid architectures: in DREMnet it operates inside disentangled denoising blocks, in RWKV-IR inside restoration-oriented transformer layers, and in AudioRWKV inside a spectrogram sequence model built around RWKV7 recurrence.

## 5. Computational complexity and implementation regimes

The central computational claim attached to Bi-WKV is that bidirectionality need not require quadratic self-attention. However, the implementation pathway matters.

For the original unidirectional WKV, DREMnet states \(O(T\!\cdot\!C)\) time and \(O(C)\) extra memory by carrying forward two \(C\)-vectors. Its Bi-WKV, written naively as a full all-to-all sum for each \(t\), has \(O(T^2\!\cdot\!C)\) time and \(O(T\!\cdot\!C)\) memory for storing \(K\) and \(V\), but the same source notes that because the decay depends only on \(|t-i|\), the two sums can be computed by two streaming passes—forward and backward—each in \(O(T\!\cdot\!C)\), yielding an optimized complexity of \(2\cdot O(T\!\cdot\!C)\) [2503.22223].

RWKV-IR states the comparison in image-restoration notation with \(T=H\cdot W\) and \(d=C\): standard dot-product attention has time \(O(T^2d)\) and memory \(O(T^2)\), Vision-RWKV with Bi-WKV on one flattening has time \(\simeq O(2Td)\) and memory \(O(Td)\), and Cross-Bi-WKV has time \(\simeq O(4Td)\) and memory \(O(Td)\). The paper further notes that Cross-Bi-WKV is approximately \(2\times\) the cost of single-direction WKV while remaining linear in \(T\) [2412.03814].

AudioRWKV gives the same asymptotic conclusion in sequence length \(L\) and channel dimension \(D\): Transformer self-attention requires \(\mathcal O(L^2D)\), causal WKV requires \(\mathcal O(LD)\), and bidirectional WKV requires two recurrences plus a fusion gate, still \(\mathcal O(LD)\). Its pseudocode is described as a minimal-cost implementation with \(\mathcal O(LD)\) time and \(\mathcal O(D)\) extra memory beyond storing inputs and outputs [2509.02167].

Taken together, these reports clarify two points often conflated. First, the bidirectional operator can be expressed in a quadratic-looking form yet implemented in linear time when the decay structure permits streaming passes. Second, “linear complexity” does not imply identical constants, buffering strategy, or memory accounting across implementations; the papers separately count recurrent state, stored \(K,V\) tensors, directional buffers, and full input-output storage.

## 6. Empirical effects, stability, and interpretive considerations

The reported empirical role of Bi-WKV is task-dependent but consistent in one respect: bidirectional context is presented as improving discrimination between signal and nuisance structure.

In DREMnet, the rationale is explicit. SATEM denoising requires each sample to depend on both past and future measurements, and Bi-WKV is said to improve the model’s ability to distinguish transient EM signal features from non-stationary noise. The learnable decay \(w\) and self-reward \(u\) are described as gates controlling how rapidly influence falls off with distance and how strongly the token admits its own information. The paper attributes improved denoising to bidirectional context, adaptive correlation-length control, and global attention applied to disentangled content and context streams, yielding higher SNR and lower MSE in both synthetic and real-field tests; it also states that processed field data more accurately reflects the theoretical signal and improves identification of subsurface electrical structures [2503.22223].

In RWKV-IR, an ablation on Urban100 \(\times2\) at 10K iterations reports PSNR values of \(32.08\) for Vision-RWKV with Q-Shift, \(32.69\) for \(+\)Q-Shift without offset shift, \(32.95\) for \(+\)DC-Shift with single Bi-WKV, and \(32.95\) for \(+\)Cross-Bi-WKV. Under Urban100 \(\times2\), 100K iterations, DF2K training, MambaIR reports \(33.55\) PSNR and \(0.9401\) SSIM, whereas RWKV-IR reports \(33.58\) PSNR and \(0.9404\) SSIM. The paper summarizes these as consistent PSNR/SSIM gains of \(0.02\)–\(0.04\) dB at linear complexity [2412.03814].

In AudioRWKV, the ablation on AudioSet-2M reports \(34.50\) mAP for causal RWKV7, \(38.39\) mAP after adding bidirectional scan with average fusion, \(39.02\) mAP after replacing average fusion by a learned gate, and \(40.91\) mAP for the final system with ConvShift and full integration. The same source states that A-RWKV-S (\(22\)M) achieves performance parity with AuM-B (\(92\)M) under the same linear-model regime, and that for long-form audio of approximately \(5\) minutes \(28\) seconds, WKV7 achieves up to a \(13.3\times\) speedup in processing; the throughput remains nearly flat while Transformer latency grows proportionally to \(L^2\), with OOM around \(L\approx2^8\) for AST(flash) [2509.02167].

AudioRWKV also attaches a specific stability argument to Bi-WKV. Because RWKV7 constrains the per-channel decay \(w_t\) to \((0,1)\), the sums \(S_t\) and \(U_t\) remain bounded, the spectral radius of the per-channel decay matrix satisfies \(\max(w_t)<1\), and Bi-WKV inherits the same numerical stability because it consists of the same bounded forward and backward recurrences followed by a convex fusion of two bounded streams. The paper reports stable convergence for models up to \(91\)M parameters, without the gradient explosions or collapses seen in large SSMs under similar regimes [2509.02167].

A useful interpretive caution follows from these results. The gains reported for “Bi-WKV” are not attributable solely to bidirectional recurrence in isolation: DREMnet combines it with disentangled content/context modeling and Covering Embedding, RWKV-IR combines it with DC-Shift and, in one variant, Cross-Bi-WKV, and AudioRWKV combines it with ConvShift and learned directional gating. The published evidence therefore supports Bi-WKV as a central component of several successful RWKV-derived systems, but not as a modality-independent drop-in explanation for the entirety of their performance.

Source: https://www.emergentmind.com/topics/bidirectional-wkv-kernel-bi-wkv