---
title: Evolving WKV Attention in DRWKV
url: https://www.emergentmind.com/topics/evolving-wkv-attention
type: topic
---

# Evolving WKV Attention in DRWKV

Searching arXiv for the cited RWKV-family papers and closely related work to ground the article in the current literature.
Evolving WKV Attention denotes a geometry-aware modification of the RWKV family’s weighted key-value accumulation in which the sequence order itself is altered to better preserve spatial structure. In the specific usage introduced by DRWKV, the mechanism replaces ordinary visual RWKV scanning with an Archimedean-spiral traversal so that receptance-gated WKV aggregation follows edge continuity rather than a generic raster order [2507.18594]. More broadly, it belongs to a larger line of RWKV research in which the original WKV recurrence is progressively extended through retrospective summaries, matrix-valued state updates, bidirectional aggregation, recurrent visual scanning, and sparse long-range retrieval, while retaining the efficiency objective that differentiates RWKV from quadratic self-attention [2411.02795].

## 1. WKV attention as the substrate of RWKV evolution

RWKV, short for Receptance Weighted Key Value, is designed to combine parallelizable training with recurrent inference by reformulating attention into a linear-time weighted key-value mechanism rather than an explicit $QK^\top$ similarity matrix [2411.02795]. In the review formulation, the parallel WKV operator over a sequence of length $T$ is

\[
\text{WKV}(K, V, W) = \frac{\sum_{i=1}^{T} \exp(k_i - (T-i)w) v_i}{\sum_{i=1}^{T} \exp(k_i - (T-i)w)}
\]

and the sequential inference form is

\[
a_t = \exp(w)a_{t-1} + \exp(k_t)
\]
\[
b_t = \exp(w)b_{t-1} + \exp(k_t)v_t
\]
\[
\text{WKV}_t = \frac{b_t}{a_t}.
\]

These equations encode a normalized exponentially decayed weighted average rather than all-pairs token interaction. The recurrent state therefore stores compressed history in constant-size accumulators, which gives constant-time update per token at inference and linear complexity in sequence length during training [2411.02795].

Receptance is the complementary gating component. In the review’s channel-mixing formulation,

\[
\text{ChannelMix}(x_t) = \sigma(W_r x_t) \odot (W_v \phi(W_k x_t)),
\]

where $\sigma$ is a sigmoid gate and $\phi$ is often $\text{ReLU}(z)^2$ [2411.02795]. This gate determines how much of the accumulated key-value information is exposed downstream. Token shifting further injects local context through

\[
\text{Shift}(x_t, x_{t-1}) = \mu x_t + (1 - \mu)x_{t-1},
\]

which gives RWKV blocks an explicitly local temporal branch alongside the recurrent WKV accumulation [2411.02795].

A useful theoretical interpretation is that WKV can be rewritten as an implicit causal self-attention operator rather than treated as a completely separate primitive. In the unified implicit-attention formulation, RWKV’s time-mixing block becomes

\[
o = \operatorname{diag}(\sigma(r))\,\hat{\alpha}\,x,
\]

where $\hat{\alpha}$ is a lower-triangular, data-dependent attention matrix induced by exponential decay, token content, token shift, and gating rather than by explicit query-key dot products [2405.16504]. This interpretation is central to later “evolutions” of WKV, because it clarifies that changing scan order, state dynamics, or auxiliary retrieval pathways changes the induced attention structure even when the model remains recurrent.

## 2. Early architectural evolutions: longer paths, richer states, stronger retrieval

Several RWKV-family papers modify WKV by changing how historical information is routed rather than by reverting to full Transformer attention. RRWKV is a direct example. It keeps RWKV’s recurrent backbone but inserts intermediate tokens called mediums, $M = \{m_1,\dots,m_c\}$, at regular intervals:

\[
X_{new} = \{ m_1, x_1, \cdots, m_c, \cdots, x_t \},
\]

with $m_1$ initialized as a zero-like sentry token [2306.05176]. Each medium summarizes a chunk of preceding tokens and is recalibrated through

\[
m_i = \sigma (W_{s_i} \cdot \delta (W_{m_i} \cdot M_i ) ).
\]

RRWKV then modifies the channel-mix pathway so that the model interpolates between the current token state and a corresponding medium:

\[
r_t = ({\nu}_r \odot o_t + (1-{\nu}_r) \odot m_t ) \cdot  W_r
\]
\[
z_t = ({\nu}_z \odot o_t + (1-{\nu}_z) \odot m_t) \cdot  W_z
\]
\[
\tilde{x}_t = \sigma (r_t) \odot (max(z_t, 0)^2 \cdot  W_v).
\]

The paper’s explicit claim is that mediums shorten the maximum path length and provide a more fluent route to distant information while keeping complexity near-linear at $O((n+c^2)\cdot d)$ [2306.05176].

Later language-model variants modify the recurrent state itself. ARWKV presents RWKV-7 as “native RWKV attention,” with a matrix-valued update

\[
State_t = State_{t-1} \left({diag}(w_{t}) - \hat{\kappa}^T_t (a_t \odot \hat{\kappa}_t)\right) + v^T_t \cdot \tilde{k}_t,
\]

contrasted with the simpler RWKV-6 recurrence

\[
state' = \text{diag}(w) \cdot state + k^\top \cdot v.
\]

The paper interprets the RWKV-7 state as a matrix-valued attention state and states that $a$ is the in-context learning rate [2501.15570]. It further claims wider eigenvalues, stronger state tracking than Transformers, and perfect 16k passkey retrieval in a 0.1B model [2501.15570]. These claims are presented as ongoing work, but they mark an important shift from scalar or vector decay toward more structured state dynamics.

RWKV-X evolves WKV in a different direction. Rather than only altering the recurrence, it keeps RWKV-7 blocks for short-range modeling and periodically inserts sparse-attention blocks for long-range retrieval [2504.21463]. The recurrent backbone is still

\[
S_t = S_{t-1}M_t + v_t^\top \cdot \tilde{k}_t,
\qquad
M_t = \text{diag}(w_t) - \hat{\kappa}_t^\top (a_t \odot \hat{\kappa}_t),
\]

but the model adds Top-$k$ Chunk Sparse Attention over chunk-pooled keys:

\[
s_i = q \cdot \left( \frac{1}{B} \sum_{j=1}^{B} k_j^{(i)} \right), \quad i = 1, \dots, n
\]
\[
\mathcal{I} = \text{TopK}\left( \{ s_i \}_{i=1}^{n}, k \right)
\]
\[
\text{Attn}(q, K_{\mathcal{I}}, V_{\mathcal{I}}) = \text{softmax}\left( \frac{q K_{\mathcal{I}}^\top}{\sqrt{d_k}} \right) V_{\mathcal{I}}.
\]

RWKV-X therefore supplements recurrent compression with explicit content-based retrieval, while claiming training complexity $O(kBN + N)$ and decoding complexity $O(1)$ [2504.21463].

| Variant | Mechanism added to WKV/RWKV | Stated objective |
|---|---|---|
| RRWKV | Medium tokens with squeeze and recalibration | Shorten path length and improve long-range dependency capture |
| ARWKV / RWKV-7 | Matrix-valued recurrent state with correction term | Increase expressiveness and state tracking |
| RWKV-X | Periodically inserted Top-$k$ chunk sparse attention | Recover explicit long-range retrieval while keeping linear/constant-time efficiency |

Taken together, these developments show that “evolution” in WKV research often means altering the information path around the recurrent accumulation rather than discarding the recurrent accumulation itself.

## 3. From 1D causality to 2D global context: bidirectional and recurrent visual WKV

Adapting WKV to images requires more than applying a text-oriented causal scan to flattened pixels or patches. The RWKV review notes that visual variants extend token shifting to 2D and introduce bidirectional attention of the form

\[
wkv_t = \frac{\sum_{i=1}^T \exp(-|t-i|w + k_i)\cdot v_i}{\sum_{i=1}^T \exp(-|t-i|w + k_i)},
\]

because spatial context is not naturally one-directional [2411.02795]. The same review also identifies Restore-RWKV as introducing Re-WKV through repeated bidirectional application,

\[
wkv_t^{(m)} = \text{Bi-WKV}(K, V^{(m-1)}),
\]

alongside Omni-Shift for 2D neighborhood aggregation [2411.02795]. This establishes a recurrent bidirectional template for later visual RWKV variants.

StyleRWKV develops this line explicitly for image style transfer. Its central modification, Re-WKV, is described as a recurrently applied bidirectional attention operator embedded in Skip Scanning, with base bidirectional aggregation

\[
{wkv_m^n} = \; \text{Bi-WKV}({K_{m}^n}_s, {V_{m}^n}_s) = \; \frac{\sum\limits_{i=1, i \neq t}^T e^{-(|t-i|-1)/T \cdot w + k_i} v_i + e^{u + k_t} v_t}{ \sum\limits_{i=1, i \neq t}^T e^{-(|t-i|-1)/T \cdot w + k_i} + e^{u + k_t} }
\]

and recurrent update

\[
{wkv_m^n}^{(j)} = \; \text{Bi-WKV}^{(j)}({K_{m}^n}_s, {wkv_m^n}^{(j-1)}),
\qquad
wkv_m^n = {wkv_m^n}^{(q)},
\]

with $q \ll T$ [2412.19535]. The stated purpose is to create a global receptive field on images while preserving linear time and memory.

StyleRWKV couples Re-WKV to S-Scanning, which reorganizes image patches using step size $p$ and an atrous-like traversal:

\[
O_i \xleftarrow{\text{scan}} {K_m^n}_s[:, a :: p, \, b :: p], \quad {K_m^n}_s^{\prime} \xleftarrow{\text{merge}} {O}_i
\]

with

\[
(a, b) = \left( \frac{1}{2} + \frac{1}{2} \sin  \left( \frac{\pi}{2} (i-2) \right), \frac{1}{2} + \frac{1}{2} \cos \left( \frac{\pi}{2} (i-2) \right) \right).
\]

The paper reports an ablation on recurrence number: $q=1$ gives ArtFID 27.783, FID 18.128, LPIPS 0.593, Time 0.205 s; $q=2$ gives ArtFID 26.370, FID 16.362, LPIPS 0.451, Time 0.213 s; and $q=3$ gives ArtFID 25.639, FID 15.442, LPIPS 0.448, Time 0.236 s [2412.19535]. The paper therefore selects $q=2$ as the best trade-off.

These visual adaptations show a recurring design pattern. WKV is preserved as a recurrent weighted key-value operator, but the sequence order, directionality, and recurrence schedule are redesigned so that the induced attention better matches 2D structure. Evolving WKV Attention in DRWKV extends exactly this pattern, but with an edge-focused topology rather than the global stylization objective of StyleRWKV.

## 4. Evolving WKV Attention in DRWKV

In DRWKV, Evolving WKV Attention is the spatial modeling mechanism of the Deep Detail Mining stage, introduced to preserve object boundaries, continuous edges, and fine structures under severe low-light degradation [2507.18594]. The paper argues that prior VRWKV-style architectures suffer from mismatch with hierarchical edge features and poor geometric modeling of edges, because edges are not well represented as uniformly distributed Euclidean features when low-light contours are better viewed as curved manifolds in a Riemannian space [2507.18594].

The mechanism is built on Evolving Scanning, a spiral sequence construction defined by the Archimedean spiral

\[
P(\theta)=\left(r(\theta)\cos\theta,\; r(\theta)\sin\theta\right),
\qquad
r(\theta)=a+b\theta.
\]

Here $a$ is the initial radius, $b$ is the spiral expansion rate, and $\theta$ is the angular parameter [2507.18594]. The scan expands gradually from local to global regions. The paper further extends this to a four-directional spiral system starting from the four corners of the feature map and supporting both clockwise and counterclockwise traversal, with the stated purpose of reducing directional bias and adapting to different edge layouts [2507.18594].

A topology-preserving constraint formalizes the continuity objective. For any two edge points $e_i, e_j \in E$, if

\[
\|e_i-e_j\|_{\text{geo}} < \delta,
\]

then their 1D sequence positions under spiral traversal should satisfy

\[
|t_i-t_j| < \tau.
\]

The intended meaning is that nearby points on the same edge should remain nearby after serialization, so that RWKV’s sequential accumulation can follow contour continuity rather than destroy it through an unsuitable scan order [2507.18594].

Within the ES-RWKV block, the spatial mix first applies Q-Shift to produce receptance, key, and value projections:

\[
(*)_s = Q\text{-}Shift_{(*)}(X)W_{(*)} = \left(X + (1-\mu_{(*)})X^\dagger\right)W_{(*)},
\qquad (*) \in \{R,K,V\},
\]

with channel-partitioned shifted feature

\[
X^{\dagger}[h,w]=Concat\bigl( X[h-1,w,0:C/4],\; X[h+1,w,C/4:C/2],\; X[h,w-1,C/2:3C/4],\; X[h,w+1,3C/4:C] \bigr).
\]

The output of the evolving spatial mix is then

\[
O_s = X + LN\big((\sigma(R_s)\odot wkv)W_{O_s}\big),
\qquad
wkv = EV\text{-}WKV(K_s,V_s).
\]

This retains the RWKV principle of receptance-gated weighted key-value accumulation, but the accumulation is performed over the spiral sequence rather than a conventional linear order [2507.18594]. In that narrow technical sense, Evolving WKV Attention is not a new all-pairs attention primitive; it is an evolution of WKV in which the scan topology is redesigned to better align recurrent aggregation with edge geometry.

## 5. Position within the DRWKV pipeline and reported empirical effects

DRWKV organizes its method around three components: Global Edge Retinex (GER), Evolving WKV Attention, and Bilateral Spectrum Aligner with MS²-Loss [2507.18594]. GER provides the decomposition model

\[
I = (R + \alpha\cdot E)\odot L + \beta\cdot N + \gamma\cdot S
\]

with refined reflectance

\[
R_{\text{enh}} = R + \alpha\cdot E.
\]

Here $R$ is reflectance, $E$ is the edge feature term, $L$ is illumination, $N$ is spatially heterogeneous noise, $S$ is artifact residual, and $\alpha,\beta,\gamma$ are weights [2507.18594]. Evolving WKV Attention is the spatial operator that the paper uses to extract and preserve the edge term $E$ during detail mining.

The overall architecture has a Light Preprocessing stage and a Deep Detail Mining stage. Light Preprocessing estimates noise, illumination, and reflectance through

\[
\hat{R}=\frac{I-N}{L}, \qquad \hat{I}=I\odot\hat{R},
\]

while Deep Detail Mining contains ES-RWKV blocks, channel mix, Bi-SAB, and SEE [2507.18594]. The paper states that Evolving WKV is used “for the first time” with wavelet downsampling to extract edge gradient features in this stage [2507.18594].

The reported computational profile is lightweight by the standards discussed in the paper. DRWKV uses only 1.67 GFLOPs in the main benchmark table, and 8.28M parameters in the ablation table [2507.18594]. The paper attributes the efficiency to retaining RWKV’s linear-time, recurrent-style computation while improving the scan order for edge structures.

The ablation evidence is reported at the level of ES-RWKV rather than a fully isolated EV-WKV row. On LOLv2-Real, the baseline gives SSIM = 0.415 and PSNR = 12.57 dB; adding ES-RWKV gives SSIM = 0.591 and PSNR = 16.27 dB; further additions eventually yield SSIM = 0.832 and PSNR = 24.12 dB [2507.18594]. The paper explicitly states that after embedding Evolving Scanning / ES-RWKV, SSIM improves by 42.4% and PSNR improves by 29.4% [2507.18594]. Qualitatively, the paper reports more continuous object edges, more stable contours, better preservation of fine details, and less blurring and fewer false edge breaks.

These results do not separate the spiral sequence from every other component in the block, but they do provide direct support for the integrated ES-RWKV design. A plausible implication is that the paper’s main empirical claim concerns the compatibility of spiral scanning with edge topology rather than a claim that scan order alone explains all of DRWKV’s gains.

## 6. Interpretation, misconceptions, and relation to adjacent RWKV research

A common misconception is that Evolving WKV Attention replaces RWKV with Transformer-style self-attention. The available formulations do not support that interpretation. In DRWKV, the aggregation remains receptance-gated weighted key-value accumulation,

\[
O_s = X + LN\big((\sigma(R_s)\odot EV\text{-}WKV(K_s,V_s))W_{O_s}\big),
\]

and the paper explicitly contrasts this with quadratic self-attention

\[
\text{Attention}(Q,K,V)=\text{softmax}(QK^\top)V
\]

to emphasize lower computational cost [2507.18594]. The modification is in sequence topology, not a return to dense all-pairs similarity.

Another misconception is that “evolving” refers only to a better recurrence equation. Across the RWKV literature, evolution has taken several distinct forms. RRWKV evolves context routing through inserted mediums [2306.05176]; ARWKV evolves the recurrent state into a more expressive matrix-valued attention memory [2501.15570]; StyleRWKV evolves WKV into recurrent bidirectional image aggregation under structured scan paths [2412.19535]; RWKV-X evolves the architecture by adding sparse long-range retrieval blocks [2504.21463]. Evolving WKV Attention in DRWKV belongs to this family of modifications, but its specific novelty is geometric serialization of edge structure rather than sparse retrieval or teacher-distilled state tracking.

The implicit-attention perspective provides a useful unifying lens. If RWKV-style models can be rewritten as $Y = \alpha(X)X$ with a data-controlled lower-triangular operator, then changing scan order, introducing bidirectionality, or adding chunk selection effectively changes the induced attention structure without abandoning recurrence [2405.16504]. This suggests that WKV evolution can be understood as the progressive design of better implicit attention operators under linear or near-linear constraints.

The broader RWKV review also frames several open problems that remain relevant here: theoretical understanding is incomplete; scaling limits are unclear; domain-specific tuning is needed; and robustness, interpretability, and hardware support remain open research areas [2411.02795]. For Evolving WKV Attention specifically, this suggests two immediate questions. First, whether topology-aware scan design can be formalized beyond the heuristic constraints already given. Second, whether edge-aware serialization generalizes to other visual tasks involving irregular structures, such as restoration or segmentation. The existing papers do not answer these questions directly, but they place Evolving WKV Attention within an increasingly diverse family of recurrent implicit-attention mechanisms that trade quadratic connectivity for carefully engineered memory paths.

Source: https://www.emergentmind.com/topics/evolving-wkv-attention