PointDGRWKV: RWKV for Domain-Generalized Point Clouds
- PointDGRWKV is a framework that re-engineers RWKV to address spatial distortions and attention drift in domain-generalized point cloud classification.
- The AGT-Shift module uses adaptive geometric token shifts via spatial hashing to capture local context efficiently without KNN computations.
- The CD-KDA module aligns key feature distributions across source domains, stabilizing global attention and boosting classification accuracy.
Searching arXiv for the target paper and closely related RWKV point-cloud work to ground the article in current literature. arxiv_search(query="(Yang et al., 28 Aug 2025)", max_results=5, sort_by="submittedDate") arxiv_search(query="(Liu et al., 9 Jun 2026)", max_results=5, sort_by="submittedDate") arxiv_search(query="PointDGMamba point cloud domain generalization", max_results=10, sort_by="relevance") PointDGRWKV is an RWKV-based framework for domain-generalized point cloud classification (DG PCC) that is designed to preserve RWKV’s linear computational complexity, global receptive field, and long-range dependency modeling while addressing two failure modes that arise when RWKV is directly transferred to unstructured 3D point clouds: spatial distortions from fixed-direction token shifts and attention drift caused by cross-domain shifts in key-feature distributions (Yang et al., 28 Aug 2025). The method introduces two task-specific modules—Adaptive Geometric Token Shift (AGT-Shift) and Cross-Domain Key Feature Distribution Alignment (CD-KDA)—and evaluates them in the setting where training uses only multiple labeled source domains, testing is conducted on a completely unseen target domain, and no target-domain data is available during training or adaptation (Yang et al., 28 Aug 2025).
1. Problem setting and motivation
Point cloud classification (PCC) assigns a semantic category to a 3D point cloud. In the standard setting, train and test data are assumed to come from the same distribution. PointDGRWKV is formulated for the domain generalization variant of PCC, where point clouds from different sensors, scanning angles, environments, occlusions, and completion pipelines exhibit substantial domain shifts, and the objective is to learn representations that remain robust on an unseen target domain without target-time adaptation (Yang et al., 28 Aug 2025).
Within this setting, RWKV is treated as an attractive backbone because it combines global receptive fields, long-range dependency modeling, and linear computational complexity. The paper positions RWKV against three commonly used architectural families in DG PCC: CNNs, which are efficient but have limited receptive fields; Transformers, which provide strong global modeling at quadratic cost; and Mamba/SSM models, whose linear efficiency is accompanied by a fixed state size that can limit long-range modeling in long point sequences (Yang et al., 28 Aug 2025). This motivates the central question addressed by PointDGRWKV: whether RWKV can be adapted to DG point cloud classification without sacrificing its linear-time appeal.
The answer given is explicitly negative for naïve transfer. The framework is therefore not presented as a direct application of an existing sequence model to 3D points, but as a re-engineering of RWKV for point-cloud geometry and cross-domain robustness (Yang et al., 28 Aug 2025).
2. Failure modes of naïve RWKV on DG point clouds
The paper identifies two core problems that arise when RWKV-style mechanisms are transplanted from sequences or images to point clouds. The first is spatial distortion induced by fixed-direction token shifts such as Q-Shift. RWKV-like vision models assume a regular layout, so shifted channel slices correspond to meaningful neighboring locations on an image grid. Point clouds are unstructured and permutation-sensitive, and fixed directional shifts therefore do not correspond to meaningful geometry. The reported consequences are distorted local spatial relationships, weakened local geometric modeling, and reduced robustness under domain shifts (Yang et al., 28 Aug 2025).
The second problem is attention drift caused by Bi-WKV’s exponential weighting over key features. The bidirectional WKV computation is described as
where and are the key and value of token , and and are learnable distance-decay and bias terms (Yang et al., 28 Aug 2025). Because keys appear inside the exponential, slight cross-domain differences in key distributions are magnified. The paper characterizes the result as attention focus drifting across domains, leading to unstable attention and degraded generalization.
These two observations define the paper’s design space. One module must replace geometry-inappropriate fixed token shifts with a point-cloud-aware alternative, and another must reduce the sensitivity of Bi-WKV attention to cross-domain variation in key statistics (Yang et al., 28 Aug 2025).
3. Model architecture and training formulation
PointDGRWKV uses a four-stage hierarchical structure with RWKV blocks stacked per stage in counts $1, 1, 2, 2$ (Yang et al., 28 Aug 2025). Training uses point-cloud preprocessing that includes scaling, normalization, and random jitter. At inference, only the trained feature extractor and classifier are used; no target-domain data is required, and no adaptation is performed, preserving compatibility with the domain generalization protocol (Yang et al., 28 Aug 2025).
The model combines a classification loss with the CD-KDA alignment loss, with default weights and (Yang et al., 28 Aug 2025). This training design places the domain-robustness mechanism entirely in the source-domain optimization stage rather than in any target-aware post hoc adjustment.
A useful way to interpret the architecture is as a decomposition of DG PCC into two coupled subproblems. The first is geometry-sensitive local context formation, handled by AGT-Shift. The second is stabilization of global RWKV attention under source-domain variation, handled by CD-KDA. This suggests that the framework treats spatial modeling and cross-domain attention stability as distinct but complementary requirements.
4. Adaptive Geometric Token Shift
AGT-Shift is introduced to replace image-style token shift with a geometry-aware local mixing mechanism that preserves efficiency (Yang et al., 28 Aug 2025). Its inputs are point features and point coordinates 0, where 1 is batch size, 2 is the number of points, and 3 is the feature dimension.
The first step is spatial partitioning through fixed step-size spatial hashing. The 3D space is discretized into grid cells 4, and points falling in the same grid cell form a local context block. The paper emphasizes that this approximates a neighborhood without explicit KNN search, pairwise distance matrices, or graph construction (Yang et al., 28 Aug 2025).
For each point 5 in grid cell 6, AGT-Shift computes a weighted local aggregation
7
with weights
8
This gives larger weights to points closer to the region center and is described as capturing local geometric context in a soft, adaptive way (Yang et al., 28 Aug 2025).
To avoid excessive disturbance of the original representation, AGT-Shift perturbs only a subset of channels. The first 9 channels are blended with the aggregated local feature using a coefficient 0, while the remaining channels are kept unchanged, and the two parts are concatenated (Yang et al., 28 Aug 2025). The module is therefore residual, selective, and geometry-aware.
The efficiency claim is explicit: AGT-Shift uses no KNN search, no explicit adjacency graph, no additional learnable parameters, and is implemented via tensor operations, yielding complexity 1 (Yang et al., 28 Aug 2025). In the paper’s framing, this is the key reason the method can improve local structure modeling without giving up RWKV’s linear efficiency.
5. Cross-Domain Key Feature Distribution Alignment
CD-KDA addresses the second failure mode: attention drift produced by domain-dependent variation in key distributions (Yang et al., 28 Aug 2025). The paper argues that the key vectors 2 differ across source domains in mean and variance/covariance, and because keys directly determine the exponential weights in Bi-WKV, even modest distribution shifts can strongly alter attention.
The alignment loss is defined over source domains 3, where the key features from domain 4 are 5. CD-KDA aligns first- and second-order statistics across unordered source-domain pairs:
6
where 7 is the domain mean of keys and 8 is the covariance matrix (Yang et al., 28 Aug 2025).
A central methodological point is that the paper prioritizes key alignment over value alignment. The stated rationale is that keys determine attention weights, whereas values are aggregated content and do not directly affect weight generation. The shared RWKV parameters 9 and 0 are also not aligned, because they are treated as inductive biases learned jointly across domains (Yang et al., 28 Aug 2025).
Conceptually, CD-KDA reduces domain-specific key shift, stabilizes the exponential weighting in Bi-WKV, and makes attention patterns more consistent across source domains. A plausible implication is that the method functions as a distributional regularizer on the attention generator rather than on the full latent representation.
6. Empirical performance, efficiency, and relation to broader RWKV-based point cloud research
The experimental protocol evaluates overall classification accuracy on held-out target domains. Two benchmarks are used: PointDA-10, with three DG settings 1, 2, and 3; and PointDG-3to1, with leave-one-out settings 4, 5, 6, and 7 (Yang et al., 28 Aug 2025). Baselines span CNN-based, Transformer-based, Mamba-based, and RWKV-based models, including PointDAN, DefRec, GAST, PDG, MetaSets, PointNeXt, X-3D, PCT, GBNet, SUG, PCM, PointDGMamba, and V-RWKV; PointRWKV is excluded because the training code was unavailable (Yang et al., 28 Aug 2025).
On PointDA-10, PointDGRWKV reports accuracies of 84.39, 54.10, and 88.49, with an average of 75.66%, exceeding PointDGMamba’s 74.85% average by +0.81 points and V-RWKV’s 72.24% average by a larger margin (Yang et al., 28 Aug 2025). On PointDG-3to1, it reports 76.37, 95.99, 63.92, and 91.38, averaging 81.92%, which is +1.39 points above PointDGMamba’s 80.53% (Yang et al., 28 Aug 2025).
| Benchmark | PointDGRWKV average | Best prior average named in the paper |
|---|---|---|
| PointDA-10 | 75.66% | 74.85% (PointDGMamba) |
| PointDG-3to1 | 81.92% | 80.53% (PointDGMamba) |
Ablation studies attribute the gain to both proposed modules. On PointDA-10, the baseline V-RWKV’ gives 72.24, adding AGT-Shift yields 73.57, adding CD-KDA yields 74.28, and using both yields 75.66 (Yang et al., 28 Aug 2025). In shifting-strategy comparisons, AGT-Shift outperforms KNN-RandOne (74.54), KNN-Avg (74.25), and KNN-WAvg (74.82) while avoiding KNN’s quadratic cost (Yang et al., 28 Aug 2025). In key/value alignment ablations, the paper reports 73.57 for none, 74.08 for only 8, 75.68 for 9 and 0, and 75.66 for only 1, and interprets this as evidence that key alignment is the critical factor (Yang et al., 28 Aug 2025).
Model scaling results are also reported. On a single RTX 4090, Ours-Base has 2.13M params, 3.22 GFLOPs, and 1.68 ms; Ours-Standard has 3.72M params, 4.57 GFLOPs, and 2.39 ms; and Ours-Large has 10.40M params, 7.60 GFLOPs, and 2.92 ms. Their corresponding PointDA-10 averages are 75.15, 75.66, and 76.13 (Yang et al., 28 Aug 2025). This supports the paper’s claim that the framework remains efficient while improving DG robustness.
The broader significance of PointDGRWKV becomes clearer when read alongside later RWKV-based point cloud work that targets a different problem formulation. “Efficient RWKV-based Representation Learning for 3D Point Clouds” introduces P-RWKV and PointER for self-supervised point cloud representation learning, emphasizing Local Perception Expansion, Spatial Context Enhancement, and bidirectional global context modeling under linear complexity (Liu et al., 9 Jun 2026). PointDGRWKV addresses domain-generalized supervised classification rather than masked autoencoding, but both works share the premise that vanilla RWKV is not well matched to irregular 3D geometry and must be adapted through explicit geometry-aware mechanisms (Liu et al., 9 Jun 2026). This suggests an emerging line of research in which RWKV serves as the efficient global backbone, while task-specific modules compensate for the mismatch between serialized sequence operations and unordered point-cloud structure.