Papers
Topics
Authors
Recent
Search
2000 character limit reached

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

Published 21 Sep 2026 in cs.LG | (2609.24797v1)

Abstract: Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in [−1,1][-1,1] and the delta-rule coefficient ββ in [0,2][0,2]. We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct2_2. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of SO(3)\mathrm{SO}(3), and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on S3S_3, S4S_4, and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.

Summary

  • The paper introduces CKDA, extending the gate range to $[-1,1] and $eta$ to $[0,2]$ within single-layer Kimi Delta Attention.
  • CKDA dynamically creates two reflections, one reflective gating and $eta = 2$ to allow a single-layer KDA head to simulate planar rotations and transitions as tracked by the system..
  • CKDA alters group tracking abilities for noncommutative transformations, enabling accurate anticipation of group behaviors such as finite word extras, automata simulation, symbolic redundancy and real-world examples such as language prediction and periodic audio processes.

Motivation and central contribution

“Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention” (2609.24797) studies the expressive limitations of linear recurrent models whose state transitions combine diagonal gating with a rank-one delta-rule update. The paper focuses on Kimi Delta Attention (KDA), whose transition for one head has the form

At=(I−βtktkt⊤)Diag⁡(αt),A_t = \left(I-\beta_t k_t k_t^\top\right)\operatorname{Diag}(\alpha_t),

where ktk_t is a normalized key, βt\beta_t controls the delta-rule update, and αt\alpha_t is a channel-wise gate. Standard KDA uses nonnegative gates and typically restricts βt\beta_t to [0,1][0,1]. These constraints preserve efficient chunkwise computation and non-expansive dynamics, but they also force the transition spectrum to remain real, preventing a single layer from representing persistent rotational dynamics.

The paper’s central proposal, Complex KDA (CKDA), combines two range extensions:

  • signed channel-wise gates αt∈[−1,1]d\alpha_t\in[-1,1]^d;
  • extended delta-rule coefficients βt∈[0,2]\beta_t\in[0,2].

The combination is essential. Extending only the gate does not produce the required reflection geometry, while extending only β\beta leaves the transition spectrally real. Together, the two extensions allow a single diagonal-plus-rank-one transition to compose two reflections: one supplied by a signed diagonal gate and one supplied by the delta-rule term at β=2\beta=2. Their noncommuting composition can be a planar rotation.

The resulting architecture retains KDA’s first-order real recurrence, fixed-size state, diagonal-plus-rank-one transition structure, non-expansiveness, and efficient implementation. The paper therefore distinguishes expressivity from merely increasing transition rank: CKDA attains state-tracking capabilities associated with two successive delta-rule updates without explicitly applying two rank-one updates per token.

The geometric mechanism is summarized by the planar construction below.

Figure 1

Figure 1: A signed diagonal reflection combined with a Householder reflection yields an arbitrary planar rotation within a single CKDA transition.

From scalar symmetry to signed channel-wise dynamics

The paper begins by comparing scalar-gated and channel-wise-gated delta transitions. In Gated DeltaNet, the gate is scalar, so the transition is

ktk_t0

The scalar factor commutes with the Householder-like delta update. Consequently, ktk_t1 is symmetric and has only real eigenvalues. Across multiple recurrent steps, scalar gates factor out as a global scale, so even signed scalar gates cannot provide an independently oriented transformation.

KDA’s channel-wise gate changes this algebra. For

ktk_t2

the two factors generally do not commute. However, the paper emphasizes that noncommutation alone is insufficient for complex eigenvalues. When all gate entries are nonnegative, the transition is similar to a symmetric matrix, and its spectrum remains real. A negative gate entry is required to make the diagonal factor indefinite.

In two dimensions, take a gate ktk_t3 and set ktk_t4. The resulting transition has the form

ktk_t5

where ktk_t6. Its eigenvalues become complex when

ktk_t7

This can occur only for ktk_t8. At ktk_t9, the diagonal gate is itself a coordinate reflection. Since the βt\beta_t0 delta update is a Householder reflection, the product becomes a rotation through angle βt\beta_t1. Varying the key direction therefore gives any planar rotation while maintaining unit spectral norm.

The paper’s stronger interpretation is that CKDA is not merely “complex” because it admits isolated complex eigenvalues. Its signed gate supplies a coordinate reflection that breaks the symmetry preventing rotational dynamics in ordinary KDA. The complex-conjugate eigenvalues are the spectral signature of this reflection composition.

Characterization of CKDA transitions

A principal theoretical result establishes that CKDA exactly contains the orthogonal diagonal-plus-rank-one family. Every orthogonal matrix of the form

βt\beta_t2

can be represented as

βt\beta_t3

where βt\beta_t4 is a diagonal sign matrix and βt\beta_t5. This is a signed Householder matrix and is a CKDA transition with βt\beta_t6 and gate entries in βt\beta_t7.

This characterization has two implications. First, CKDA captures all orthogonal transformations available within the diagonal-plus-rank-one class; adding an asymmetric rank-one correction does not enlarge the orthogonal family in the relevant sense. Second, the rank-one structure imposes a sharp spectral restriction: a single non-expansive CKDA transition can contain at most one non-real conjugate eigenvalue pair on the unit circle.

More generally, the paper decomposes a CKDA transition into a sign-magnitude component and a signed Householder component. The latter acts nontrivially in at most a two-dimensional subspace determined by the key’s positive and negative gate components. The remaining coordinates receive only real eigenvalues. Complex eigenvalues require all of the following:

  1. at least one negative gate entry;
  2. a key with support across coordinates having opposite gate signs;
  3. βt\beta_t8.

At βt\beta_t9, the active two-dimensional block is an exact rotation. For intermediate αt\alpha_t0, it is a non-expansive contraction-rotation whose eigenvalues can move continuously from the real axis toward the unit circle.

The paper also proves factorization results for products of CKDA transitions. Any non-expansive αt\alpha_t1 matrix can be represented using at most αt\alpha_t2 CKDA factors, while any orthogonal matrix requires at most αt\alpha_t3 factors. The latter bound is sharp: an αt\alpha_t4-cycle permutation matrix cannot be expressed using fewer than αt\alpha_t5 diagonal-plus-rank-one factors. Thus CKDA does not make every matrix a single transition; rather, it maximizes the orthogonal expressivity available from one such transition and reduces the factor count for structured products.

State-tracking expressivity

The paper evaluates recurrent expressivity through state tracking, in which the model must compose input-dependent transformations over arbitrary sequence lengths. Group-word problems are particularly appropriate because they require order-sensitive, noncommutative composition rather than simple additive counting.

The main single-layer theorem states that one CKDA head can track every finite group isomorphic to a subgroup of αt\alpha_t6. Specifically, CKDA realizes:

  • all finite cyclic groups αt\alpha_t7 and dihedral groups αt\alpha_t8 in two dimensions;
  • αt\alpha_t9 and βt\beta_t0 in three dimensions;
  • βt\beta_t1 in four dimensions using a many-to-one decoder.

The two-dimensional result follows directly from planar rotations and reflections. The βt\beta_t2 construction uses the rotational symmetry group of the cube. By conjugating the cube representation by a suitable global rotation, all relevant rotation axes can be placed in coordinate planes, satisfying the signed-Householder criterion.

The βt\beta_t3 result is more subtle. The icosahedral rotation group cannot be represented by signed Householder matrices in three dimensions: too many order-three and order-five rotation axes fail the coordinate-plane condition under every orientation. In four dimensions, however, the authors use the quaternionic relationship between βt\beta_t4 and βt\beta_t5 to encode each three-dimensional rotation as a signed Householder transition. A fixed decoder recovers the accumulated βt\beta_t6 rotation, although the hidden βt\beta_t7 state contains additional information. The resulting tracker has at most

βt\beta_t8

reachable hidden matrices for the βt\beta_t9 group elements.

This distinction between realization and tracking is important. A realization requires the hidden matrices themselves to form a faithful group representation. A tracker permits several hidden states to decode to the same target state. The four-dimensional [0,1][0,1]0 construction relies on the latter, and its representability does not imply easy optimization.

The paper also gives a negative result for [0,1][0,1]1. Under non-expansive transitions, finite reachability, and at most one persistent complex-conjugate pair per head, no single CKDA layer with any finite number of independent heads can track [0,1][0,1]2. The proof reduces finite affine tracking to a finite group of orthogonal transitions and constructs a commutator whose tenth power is the identity in every head, while its decoded permutation is a three-cycle whose tenth power is nonidentity. This obstruction applies despite arbitrary additive updates and many-to-one decoders.

The resulting separation is precise: CKDA handles [0,1][0,1]3 in one layer but not [0,1][0,1]4 under the stated assumptions. The obstruction is not simply dimensional; it follows from the spectral limitation of rank-one non-expansive transitions.

Multi-layer simulation of automata

The paper extends the single-layer results to general finite-state and weighted computations. Three CKDA layers suffice for every finite group-word problem. If expansive updates with [0,1][0,1]5 are permitted, three layers also suffice for every regular language and for weighted finite automata over [0,1][0,1]6 under polynomial-precision exact arithmetic.

The construction uses a clock-buffer-accumulator organization. The clock tracks position modulo a fixed period, the buffer stores a bounded block of recent inputs, and the accumulator applies a factorization of the corresponding block transition. CKDA’s planar rotation allows the clock to be implemented in one layer. This removes one clock layer relative to constructions based on scalar-gated DeltaNet, yielding a three-layer rather than four-layer simulation.

The result depends on explicit arithmetic assumptions. The constructive proofs use exact arithmetic over a fixed real algebraic number field with polynomially bounded rational coefficients. They do not establish equivalent guarantees for ordinary floating-point execution. The WFA construction also requires [0,1][0,1]7, which sacrifices the non-expansiveness guarantee central to the stable CKDA regime.

Empirical evaluation on symbolic and periodic tasks

Group-word extrapolation

The authors train one-layer models on [0,1][0,1]8, [0,1][0,1]9, and αt∈[−1,1]d\alpha_t\in[-1,1]^d0 word problems, with training lengths up to αt∈[−1,1]d\alpha_t\in[-1,1]^d1 and evaluation at longer lengths. Among the tested KDA parameter ranges, only the joint extension to signed gates and αt∈[−1,1]d\alpha_t\in[-1,1]^d2 consistently provides strong extrapolation on αt∈[−1,1]d\alpha_t\in[-1,1]^d3 and αt∈[−1,1]d\alpha_t\in[-1,1]^d4S_3αt∈[−1,1]d\alpha_t\in[-1,1]^d50.2,approximatelytheperformanceexpectedfromdistinguishingonlythetwoparity−likecosets.</p><p>ThelearnedCKDAmodelrecoversthetheoreticalmechanismratherthanmerelyexploitinganunidentifiedparameterregime.Asuccessfulheadapproaches, approximately the performance expected from distinguishing only the two parity-like cosets.</p> <p>The learned CKDA model recovers the theoretical mechanism rather than merely exploiting an unidentified parameter regime. A successful head approaches \alpha_t\in[-1,1]^d6,itsgatebecomesnearlysign−valued,anditstransitionspectrumdevelopsacomplex−conjugatepairneartheunitcircle.</p><p><imgsrc="https://images.emergentmind.com/paper−images/2609−24797/s3signedkdainterpretability.png"alt="Figure2"title=""class="markdown−image"loading="lazy"></p><p><pclass="figure−caption">Figure2:Learned6, its gate becomes nearly sign-valued, and its transition spectrum develops a complex-conjugate pair near the unit circle.</p> <p><img src="https://images.emergentmind.com/paper-images/2609-24797/s3_signed_kda_interpretability.png" alt="Figure 2" title="" class="markdown-image" loading="lazy"></p> <p><p class="figure-caption">Figure 2: Learned \alpha_t\in[-1,1]^d$7 transitions use near-reflection updates, nearly binary signed gates, low-dimensional key structure, and complex eigenvalues.

The $\alpha_t\in[-1,1]^d$8 construction was not learned from standard random initialization. It became learnable under a theory-informed quaternion initialization, with a modified architecture and training schedule. This result supports the representability theorem but simultaneously demonstrates that representational capacity and optimization accessibility are distinct.

Periodic waveform continuation

The periodic audio experiment tests whether complex modes support phase preservation beyond the training horizon. A single recurrent layer observes a half-bar cue and must continue a periodic waveform after the input becomes zero. Models are trained up to sequence length $\alpha_t\in[-1,1]^d$9 and evaluated through length $\beta_t\in[0,2]$0.

At length $\beta_t\in[0,2]1,CKDAachievesareported<ahref="https://www.emergentmind.com/topics/signal−to−noise−ratio−snr"title=""rel="nofollow"data−turbo="false"class="assistant−link"x−datax−tooltip.raw="">signal−to−noiseratio</a>of<strong>38.1dB</strong>,whereasthecausalTransformerfallsto<strong>2.8dB</strong>.TheotherKDArangeconfigurationsfailtomaintainaccuratephaseoverthesameextrapolationhorizon.Thisisdirectevidencethattherotationalmechanismlearnedinsymbolictaskscanalsosupportcontinuousperiodiccontinuation.</p><p>TheresultisnotuniformlyfavorabletoCKDA:a<ahref="https://www.emergentmind.com/topics/gated−recurrent−unit−gru−block"title=""rel="nofollow"data−turbo="false"class="assistant−link"x−datax−tooltip.raw="">GRU</a>achievesthelowestwaveformerror.Hencetheexperimentisolatesthevalueofstablelinearrotationaldynamicsforextrapolation,notoverallsuperiorityovernonlinearrecurrentmodels.</p><p><imgsrc="https://images.emergentmind.com/paper−images/2609−24797/groovewaveformmain.png"alt="Figure3"title=""class="markdown−image"loading="lazy"></p><p><pclass="figure−caption">Figure3:CKDAmaintainsperiodicwaveformphasesubstantiallybeyondthemaximumtraininglength,althoughtheGRUremainsmoreaccurate.</p></p><h2class=′paper−heading′id=′language−modeling−and−systems−results′>Languagemodelingandsystemsresults</h2><p>Thelanguage−modelingexperimentsassesswhethertheadditionalrecurrenceexpressivitytranslatesintoimprovementsonnatural−languageobjectives.Theanswerisqualified.CKDAgenerallyperformsonparwithKDA,whilebothoutperformtheTransformerandseveralrecurrentbaselinesinthereportedparameter−matchedsettings.</p><p>Inthe1, CKDA achieves a reported <a href="https://www.emergentmind.com/topics/signal-to-noise-ratio-snr" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">signal-to-noise ratio</a> of <strong>38.1 dB</strong>, whereas the causal Transformer falls to <strong>2.8 dB</strong>. The other KDA range configurations fail to maintain accurate phase over the same extrapolation horizon. This is direct evidence that the rotational mechanism learned in symbolic tasks can also support continuous periodic continuation.</p> <p>The result is not uniformly favorable to CKDA: a <a href="https://www.emergentmind.com/topics/gated-recurrent-unit-gru-block" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">GRU</a> achieves the lowest waveform error. Hence the experiment isolates the value of stable linear rotational dynamics for extrapolation, not overall superiority over nonlinear recurrent models.</p> <p><img src="https://images.emergentmind.com/paper-images/2609-24797/groove_waveform_main.png" alt="Figure 3" title="" class="markdown-image" loading="lazy"></p> <p><p class="figure-caption">Figure 3: CKDA maintains periodic waveform phase substantially beyond the maximum training length, although the GRU remains more accurate.</p></p> <h2 class='paper-heading' id='language-modeling-and-systems-results'>Language modeling and systems results</h2> <p>The language-modeling experiments assess whether the additional recurrence expressivity translates into improvements on natural-language objectives. The answer is qualified. CKDA generally performs on par with KDA, while both outperform the Transformer and several recurrent baselines in the reported parameter-matched settings.</p> <p>In the \beta_t\in[0,2]$2M-parameter Nemotron-CC experiments, the best CKDA configuration reaches 52.30% average downstream accuracy, compared with 51.32% for standard KDA. In the smaller FineWeb ablations, CKDA reaches 52.21%, compared with 51.85% for KDA without SiLU on keys. These differences are modest and configuration-dependent; CKDA does not consistently minimize validation perplexity.

At $\beta_t\in[0,2]$3B parameters trained on $\beta_t\in[0,2]$4B FineWeb-Edu tokens, the non-hybrid models report:

Model WikiText perplexity LAMBADA perplexity Average accuracy
KDA, bounded gate 15.73 10.53 54.09%
CKDA 15.78 10.08 54.06%
Gated DeltaNet-2 15.90 11.41 53.11%
Mamba-3 MIMO 16.45 11.66 52.39%
Transformer hybrid baseline 19.22 13.72 50.86%

The central language-modeling claim is therefore not that CKDA decisively improves perplexity over KDA. Rather, CKDA preserves KDA-level language performance while adding a theoretically meaningful rotational mechanism and retaining almost all of KDA’s throughput. The implementation achieves approximately 96–97% of standard KDA throughput. The authors explicitly avoid claiming an inherent computational advantage over DeltaProductβt∈[0,2]\beta_t\in[0,2]5, whose measured throughput is competitive despite applying two delta-rule updates per token.

Scaling experiments across βt∈[0,2]\beta_t\in[0,2]6M to βt∈[0,2]\beta_t\in[0,2]7B parameters and βt∈[0,2]\beta_t\in[0,2]8B to βt∈[0,2]\beta_t\in[0,2]9B tokens show that CKDA and KDA consistently outperform the Transformer baseline in the reported grid. At β\beta0B parameters and β\beta1B tokens, CKDA obtains validation loss 2.1061, compared with 2.1630 for the Transformer and 2.1046 for KDA. The difference between CKDA and KDA is negligible at this scale, supporting the paper’s more restrained conclusion that CKDA has comparable scaling behavior rather than a clear scaling-law advantage.

Hybrid models produce a similar pattern. CKDA with a β\beta2 recurrent-to-attention ratio obtains average downstream accuracy close to the corresponding KDA hybrid, with differences varying across evaluations. The experiments also show that complex transitions emerge during language-model training: negative gates appear most often in early layers, β\beta3 occurs throughout the network, and complex eigenvalues are concentrated particularly in the first two layers.

Figure 4

Figure 4: In trained β\beta4B CKDA models, negative gates, extended β\beta5 values, and complex transition spectra emerge during optimization, especially in early layers.

The mechanistic interpretation remains incomplete. The paper identifies when the extended ranges are used but does not determine which computations they implement in language modeling.

Limitations and open questions

CKDA remains constrained by its rank-one transition structure. A non-expansive transition of this type supports at most one persistent complex-conjugate eigenvalue pair, so it cannot provide a full bank of independent oscillatory modes within one head. Products of CKDA transitions are more expressive, but the factor count increases with dimension.

The theoretical results also depend on assumptions that limit their direct interpretation for practical floating-point models. The constructive state-tracking theorems use exact finite datatypes or polynomial-precision arithmetic over fixed algebraic number fields. They do not provide error bounds for approximate rotations under arbitrarily long recurrent execution. The β\beta6 impossibility theorem additionally assumes finite reachability, which is natural for the exact group constructions but does not automatically follow from finite-precision computation.

Representability does not imply learnability. Standard training failed to discover the β\beta7 solution, and the successful experiment used theory-based initialization together with a different setup. The paper therefore leaves open whether generic optimization can reliably discover the relevant signed-reflection geometry.

The language-modeling evidence also does not establish a broad empirical advantage. CKDA matches KDA more closely than it surpasses it, and the main improvements over Transformer baselines may reflect the broader recurrent architecture, training recipe, or hybrid design rather than the signed-gate mechanism itself. The reported hybrid comparisons use different attention configurations from some baselines, and the authors acknowledge that this complicates causal attribution.

Finally, the functional roles of negative gates and β\beta8 in trained LLMs remain unresolved. The paper observes their emergence but does not provide a mechanistic decomposition connecting individual complex transitions to specific long-context or linguistic computations.

Conclusion

The paper establishes that signed channel-wise gating fundamentally changes the expressive geometry of KDA. When combined with β\beta9, it lets one non-expansive diagonal-plus-rank-one transition compose two reflections and realize arbitrary planar rotations. CKDA consequently captures the full orthogonal diagonal-plus-rank-one family, supports one-layer tracking of substantial nonabelian groups including β=2\beta=20 and β=2\beta=21, and reduces the layer count of several automata constructions.

The empirical results validate the mechanism most clearly on symbolic extrapolation and periodic waveform continuation: CKDA is the only tested KDA range configuration that reliably preserves the relevant long-horizon structure, achieving 38.1 dB SNR at length 264 in the waveform task. In language modeling, its primary result is parity with KDA rather than a decisive improvement, together with near-baseline throughput and evidence that the extended parameter ranges are used after training. The remaining central question is whether the theoretical rotational capacity can be made reliably learnable and translated into a reproducible advantage on workloads where long-horizon state composition is a limiting factor.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces a new type of neural-network memory called Complex KDA, or CKDA.

Neural networks that work with sequences—such as sentences, music, or time series—need to remember what happened earlier. CKDA is designed to do this efficiently, especially when sequences are very long.

The researchers’ main idea is to improve Kimi Delta Attention (KDA) by allowing parts of its memory update to use negative values. This small-looking change lets CKDA perform something very important: rotations.

A rotation is useful because it allows a model to keep track of repeating patterns, such as:

  • counting around a circle,
  • remembering the order of objects being moved,
  • continuing a repeating musical rhythm,
  • tracking complicated combinations of actions.

2. What questions did the researchers ask?

The paper focuses on several main questions:

  1. Can KDA remember more complicated patterns without becoming much slower?
  2. Can one KDA update perform a rotation, instead of needing two separate updates?
  3. Does allowing negative “gates” make the model more powerful?
  4. Can CKDA remember information over much longer sequences than ordinary KDA?
  5. Does this extra mathematical ability improve real tasks such as language modeling and audio continuation?

Here, a gate is like a control knob. It decides how strongly different parts of the model’s memory should be kept, changed, or reversed.

3. How did the researchers study this?

The researchers used both mathematics and computer experiments.

Mathematical analysis

They examined the rule CKDA uses to update its memory. In simplified form, the update is:

new memory=delta update×channel-wise gate.\text{new memory} = \text{delta update} \times \text{channel-wise gate}.

The delta update changes the memory using information from the current input. The channel-wise gate gives each memory channel its own control value.

A useful analogy is a group of arrows on a piece of paper:

  • A normal update might stretch or shrink the arrows.
  • A reflection flips them across a line.
  • Two reflections at different angles create a rotation.

The researchers showed that CKDA can combine:

  1. a reflection controlled by the delta rule, and
  2. another reflection created by signed gates.

Together, these can produce rotations while keeping the update efficient.

They also studied the update matrices. A matrix is simply a table of numbers that describes how a model changes its memory. The researchers proved that CKDA can represent every orthogonal matrix of a certain useful form. Orthogonal means that the transformation preserves lengths, like a rotation or reflection.

State-tracking experiments

The researchers tested whether CKDA could remember the results of sequences of actions. For example, imagine a shell game with three objects. If objects are swapped many times, the model must remember where each object ended up.

They tested group-tracking problems involving:

  • S3S_3: arrangements of three objects,
  • S4S_4: arrangements of four objects,
  • A5A_5: a more difficult mathematical structure related to the symmetries of an icosahedron.

They trained the models on shorter sequences and then tested them on longer sequences. This is called length extrapolation: seeing whether a model can work beyond the sequence lengths it practiced.

Audio experiment

The researchers also gave the models part of a repeating musical waveform and asked them to continue it after the input stopped.

This tests whether the model can remember the phase of a repeating pattern—essentially, where it is in the cycle.

Language-modeling experiments

Finally, they trained large models with about 1.3 billion parameters on 100 billion tokens of educational web text. They compared CKDA with:

  • Transformers,
  • Mamba-style models,
  • Gated DeltaNet,
  • ordinary KDA,
  • other recurrent neural networks.

They measured language-model quality, reasoning-task accuracy, and computation speed.

4. What did they find?

CKDA can perform rotations in one update

The most important theoretical result is that CKDA can create a 2D rotation using only one delta-rule update.

Ordinary KDA generally uses nonnegative gates. With only nonnegative gates, its transformations have limited behavior and cannot easily create true rotations.

CKDA allows:

  • gates between −1-1 and $1$,
  • delta-rule strength β\beta between $0$ and $2$.

The negative gates can flip one part of the memory while leaving another part unchanged. Combined with another reflection, this creates a rotation.

This gives CKDA more expressive power without requiring two separate delta updates.

It can track important groups with one layer

The theory shows that a single CKDA layer can track:

  • cyclic patterns such as counting around a loop,
  • dihedral patterns involving rotations and reflections,
  • the symmetry groups A4A_4 and S4S_4,
  • and, with a slightly larger representation, A5A_5.

In simpler terms, CKDA can remember complicated order-sensitive combinations of actions using fewer layers than several competing recurrent models.

However, the researchers also proved a limit: one CKDA layer cannot track S5S_5 under their stability conditions. This is important because it shows that CKDA is more powerful, but not unlimited.

Better long-sequence state tracking

On the S3S_3 and S4S_4 tests, CKDA was much better at remembering information over long sequences than the other tested KDA settings.

Using only one of the two improvements—negative gates or the larger delta range—was not enough. The strongest results came from using both together.

The model also learned the behavior predicted by the mathematics:

  • delta strength close to β=2\beta=2,
  • gates close to either −1-1 or +1+1,
  • complex-valued eigenvalues.

An eigenvalue is a mathematical way to describe how a transformation behaves. Complex eigenvalues often indicate that part of the system is rotating rather than simply growing or shrinking.

Strong periodic-audio continuation

CKDA was especially good at continuing the repeating audio waveform.

At a test length of 264 steps—well beyond the longest training length of 136 steps—CKDA achieved a signal-to-noise ratio of 38.1 dB. The causal Transformer achieved only 2.8 dB on this test.

This suggests that rotations help the model preserve repeating patterns over long periods.

The GRU still performed better on this particular audio task, so CKDA was not the best model in every situation.

Competitive language modeling

In language modeling, CKDA performed about as well as ordinary KDA and better than the Transformer and other recurrent models included in the comparison.

For example, on the reported language benchmarks:

Model WikiText perplexity LAMBADA perplexity Average reasoning accuracy
Transformer hybrid 19.22 13.72 50.86%
Gated DeltaNet 16.40 11.89 52.07%
KDA 16.81 11.68 52.28%
CKDA 15.78 10.08 54.06%

Lower perplexity is better because it means the model is less surprised by the text. Higher reasoning accuracy is better.

CKDA also kept about 96–97% of standard KDA’s processing speed, so the additional expressive power did not cause a major slowdown.

5. Why is this important?

Many sequence models face a trade-off:

  • Simple models are fast but may struggle with complicated patterns.
  • More powerful models can remember more, but may require more computation or memory.

CKDA tries to improve this trade-off. It adds a small change—signed, channel-wise gates—but gains the ability to represent rotations and more complicated state changes.

This could be useful for:

  • long-context LLMs,
  • speech and music processing,
  • periodic signals,
  • tracking actions and object movements,
  • efficient models for devices with limited memory or computing power.

The work also gives researchers a clearer understanding of why some recurrent models are more expressive than others. It shows that negative gates are not merely a minor numerical detail: they can change the geometry of the memory updates and allow the model to represent rotations.

Still, the results should be interpreted carefully. The strongest theoretical results rely on exact or carefully controlled mathematical settings, and some experiments use special initialization or simplified tasks. CKDA also does not solve every state-tracking problem, and its language-modeling improvement over KDA is modest.

Overall, the paper argues that Complex KDA is an efficient recurrent architecture that can remember certain long-term and repeating patterns better than standard KDA, while remaining fast enough for large-scale use.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Optimization versus representational capacity remains unresolved: CKDA can represent rotations and group actions theoretically, but standard training fails on A5A_5 without theory-specific initialization; the causes of this optimization barrier and methods for overcoming it are not systematically studied.
  • The benefit of signed gates is not isolated in language modeling: CKDA performs similarly to bounded-gate KDA, but the experiments do not establish whether its additional rotational expressivity improves language modeling beyond changes in initialization, parameterization, optimization, or implementation.
  • Scaling behavior is insufficiently characterized: Language-model results are reported at only approximately $1.3$B parameters and $100$B tokens, leaving open whether CKDA retains an advantage at substantially larger model sizes, longer training budgets, and different data mixtures.
  • The comparison with Transformer and recurrent baselines is not fully controlled: The models may differ in training recipes, architectures, tokenization, normalization, parameter allocation, and implementation details, making it difficult to attribute performance differences specifically to CKDA’s transition structure.
  • Long-context language modeling is not directly evaluated: The paper provides synthetic extrapolation and standard benchmark results, but does not measure perplexity, retrieval, or reasoning performance as a function of context length substantially beyond the training context.
  • Numerical robustness of signed, near-orthogonal dynamics is unclear: Theoretical constructions use exact or fixed-precision arithmetic, whereas practical experiments use BF16; the accumulation of phase, sign, and amplitude errors over very long sequences is not quantitatively analyzed.
  • Error-correction capability is unexplored: CKDA can preserve rotations, but the paper does not determine whether it can recover from noisy, corrupted, or ambiguous state updates, which is important for robust long-horizon tracking.
  • The effect of gate magnitudes away from {±1}\{\pm1\} is underexplored: The theory emphasizes nearly sign-valued gates and β≈2\beta\approx2, but it is unclear how performance changes across the continuous parameter ranges [−1,1][-1,1] and [0,2][0,2], or whether intermediate values provide useful stability–adaptivity trade-offs.
  • The interaction between non-expansiveness and practical forgetting is unresolved: Orthogonal transitions preserve information but do not inherently provide selective forgetting; the paper does not characterize how CKDA balances persistent memory with rapid state reset in realistic tasks.
  • The claimed single-layer expressivity bounds are not known to be tight: The paper proves that one layer can track several groups and that one-layer S5S_5 tracking is impossible under specified assumptions, but the exact boundary for other finite groups, especially SnS_n for n>5n>5, remains open.
  • The S5S_5 impossibility result depends on restrictive assumptions: It assumes non-expansive transitions, finite reachability, and a finite number of independent heads; it remains unknown whether one-layer CKDA can track S5S_5 with expansive transitions, infinite or approximate reachability, shared cross-head state, or nonstandard decoders.
  • Approximate state tracking is not theoretically characterized: The impossibility results concern exact decoding and finite reachability, whereas practical models operate with approximate representations; robustness thresholds and accuracy–length trade-offs are not derived.
  • The role of multiple heads is incompletely understood: The paper rules out S5S_5 under independent-head assumptions but does not determine how shared parameters, cross-head mixing, head-specific dimensions, or nonlinear inter-head decoders alter expressivity.
  • The minimal dimension and layer count are unknown: The constructions establish sufficient dimensions and depths for several groups and automata, but do not generally provide minimal state dimension, number of heads, or number of CKDA layers.
  • The relationship between CKDA products and minimal factorization complexity is incomplete: Upper bounds on the number of CKDA factors are given, but the exact minimum number of factors required for important classes of orthogonal or non-expansive matrices is not established beyond the stated sharp orthogonal bound.
  • The scope of the one-complex-pair limitation is unclear for learned products: Each CKDA transition has at most one non-real conjugate eigenvalue pair, while products can be universal, but the paper does not characterize how many steps, layers, or factors are needed to generate specified multi-frequency dynamics.
  • Generalization beyond finite groups is limited: The empirical state-tracking evaluation focuses mainly on S3S_3, S4S_4, and A5A_5; performance on non-group finite automata, semigroups, modular arithmetic with noisy inputs, hierarchical counters, and compositional algorithmic tasks remains untested.
  • Theoretical constructions for regular languages and weighted finite automata are not empirically validated: The paper proves multilayer results, including constructions requiring β>2\beta>2, but does not demonstrate these capabilities on representative regular-language or WFA computation benchmarks.
  • The practical consequences of allowing β>2\beta>2 are unresolved: Expansive transitions are required for some theoretical constructions, yet their numerical stability, training dynamics, precision requirements, and effect on long-context language modeling are not experimentally assessed.
  • The audio experiment has limited breadth: It uses a single synthetic periodic groove and a single-forward-pass continuation setup; it does not establish whether CKDA benefits real music, speech, multiscale periodic signals, irregular frequencies, changing tempo, or noisy observations.
  • The advantage over nonlinear recurrent models is not established: The GRU outperforms CKDA on the reported audio task, but the paper does not investigate whether hybrid CKDA–nonlinear architectures can combine rotational extrapolation with nonlinear phase inference.
  • The effect of input and output nonlinearities is insufficiently separated: State-tracking experiments remove the inherited SiLU activation, whereas language-model experiments ablate it only partially; the role of key/query nonlinearities in enabling or suppressing rotational dynamics remains unclear.
  • Initialization sensitivity is not systematically measured: The paper uses spread signed-gate initialization and specialized initialization for A5A_5, but does not report success rates across seeds, initialization scales, gate-sign proportions, or training schedules.
  • Throughput and memory claims lack broader systems evaluation: The reported $96$–97%97\% of KDA throughput is measured for a specific H100 configuration and sequence shape; scaling across GPUs, sequence lengths, batch sizes, head dimensions, backward passes, memory usage, and distributed training is not established.
  • The cost of signed-gate handling during inference is not fully quantified: The paper discusses gauge transformations and fused kernels, but does not provide end-to-end latency, energy, memory-bandwidth, or deployment costs relative to KDA and DeltaProduct2_2.
  • The stability analysis is primarily spectral and structural: Non-expansive transitions are shown to preserve norm bounds, but the paper does not analyze transient amplification in non-normal products, gradient stability, sensitivity to perturbations, or stability of the complete nonlinear input-dependent network.
  • The decoder’s role may overstate practical tracking ability: Several theoretical results permit sufficiently expressive feed-forward or many-to-one decoders; the complexity, parameter count, training feasibility, and generalization of these decoders are not bounded.
  • The relationship between learned CKDA transitions and the prescribed group representations is only qualitatively examined: The interpretability analysis shows complex spectra and near-reflection parameters for S3S_3, but does not quantify representation alignment, transition error, subgroup closure, or how consistently the learned mechanism matches the theoretical construction.
  • No systematic robustness evaluation is provided: Effects of token noise, missing inputs, adversarial perturbations, quantization, reduced precision, sequence truncation, and distribution shift on state tracking and language modeling remain unknown.
  • The usefulness of CKDA for tasks requiring non-periodic and content-dependent memory is uncertain: The strongest evidence concerns rotations and periodic continuation, so it remains unclear whether signed rotational dynamics improve retrieval, induction, copying, hierarchical reasoning, or other non-periodic sequence behaviors.

Practical Applications

Immediate Applications

  • Deploy CKDA as a long-context language-model backbone (software, NLP, enterprise AI). The released implementation and models can be used to build recurrent LLMs for document processing, retrieval-augmented generation, code completion, and streaming text generation. CKDA is particularly relevant where linear-time sequence processing and fixed-size recurrent state are preferable to quadratic self-attention.
    • Practical workflow: replace or augment KDA/DeltaNet blocks in an existing decoder-only LLM, fine-tune on domain data, and benchmark perplexity, throughput, memory use, and long-context accuracy.
    • Evidence from the paper: CKDA retains approximately 96–97% of standard KDA throughput and performs comparably to KDA while outperforming the listed Transformer and linear-RNN baselines in the reported 1.3B-parameter experiment.
    • Dependencies: results depend on model scale, training data, initialization, hardware kernels, numerical precision, and whether signed gates are successfully optimized. The reported results do not establish superiority across all tasks or scales.
  • Efficient streaming inference for long sequences (edge computing, communications, monitoring, embedded AI). CKDA’s fixed-size recurrent state and linear sequence scaling can support online processing of sensor streams, logs, transcripts, and user interactions without storing the entire history.
    • Potential products: streaming speech or text assistants, online log summarizers, anomaly-monitoring agents, and low-memory document or message processors.
    • Why it is actionable: the recurrence is compatible with existing KDA-style kernels and does not require an auxiliary recurrent state or explicit phase representation.
    • Dependencies: recurrent-state compression may lose information on tasks requiring arbitrary retrieval; production systems should compare CKDA against attention or hybrid architectures for factual recall and error accumulation.
  • Periodic signal continuation and phase tracking (audio, industrial monitoring, robotics, control). CKDA can be used for extrapolating periodic or quasi-periodic signals after an observed prefix, such as audio waveforms, vibration signals, rotating-machine telemetry, or cyclic actuator trajectories.
    • Potential workflow: infer the phase and latent state from a short cue, then generate or forecast future values in a single forward pass.
    • Evidence from the paper: CKDA achieved 38.1 dB SNR at a continuation length beyond the training horizon on the synthetic periodic-groove task, whereas the causal Transformer degraded substantially.
    • Dependencies: the experiment is synthetic and periodic; performance on noisy, drifting, multi-frequency, or nonstationary real-world signals remains unverified. A GRU was more accurate on the reported task, so CKDA should be treated as an efficient alternative rather than a universally superior forecaster.
  • Long-horizon symbolic and permutation-state tracking (education technology, software testing, algorithmic reasoning). CKDA can serve as a compact recurrent model for tasks involving parity, modular counting, cyclic transformations, and permutation composition.
    • Potential tools: sequence-learning benchmarks, differentiable finite-state machines, automated reasoning modules, and models that track object swaps or action sequences.
    • Evidence from the paper: a single CKDA layer tracks finite cyclic and dihedral groups and successfully extrapolates on the reported S_3 and S_4 tasks.
    • Dependencies: exact or sufficiently stable numerical representations, a suitable decoder, and training procedures that discover the rotational mechanism are important. These benchmark capabilities do not automatically imply robust general reasoning in unconstrained environments.
  • Use as an experimental drop-in for existing KDA systems (academic and industrial ML engineering). The open-source implementation can support controlled ablations of gate ranges and delta-rule coefficients.
    • Actionable comparison: evaluate standard KDA, bounded positive-gate KDA, signed-gate KDA, and CKDA on the same data, measuring extrapolation, throughput, memory, numerical stability, and optimization behavior.
    • Dependencies: signed-gate implementations require correct sign handling in forward and backward passes; kernel support and hardware compatibility may affect the claimed throughput.
  • Compact recurrent modeling for on-device applications (mobile, IoT, robotics). Because CKDA uses diagonal-plus-rank-one transitions and non-expansive updates, it is a candidate for memory-constrained inference in devices processing continuous input.
    • Potential applications: wearable-sensor interpretation, robot activity recognition, predictive maintenance, and local voice or gesture processing.
    • Dependencies: the paper evaluates GPU kernels rather than complete low-power deployments. Quantization, power consumption, latency under batch size one, and robustness to sensor noise require separate validation.

Long-Term Applications

  • Long-context foundation models with recurrent or hybrid architectures (NLP, multimodal AI). CKDA could become a recurrent alternative to attention in models that need to process very long documents, code repositories, genomic sequences, or multimodal event streams.
    • Potential architecture: use CKDA for the bulk of sequence processing and reserve attention or external memory for retrieval-heavy operations.
    • Expected benefit: learned rotational dynamics may preserve phase, order, and periodic structure over longer horizons while keeping linear inference cost.
    • Dependencies: scaling behavior beyond the reported 1.3B-parameter experiment must be established. Important unresolved issues include optimization at large scale, recurrent-state corruption, token-level retrieval, parallel training efficiency, and performance on real long-context benchmarks.
  • Robust finite-state and automaton emulation (formal methods, verification, program synthesis). The paper’s group-tracking results suggest using CKDA as a differentiable implementation of finite-state machines, regular languages, and—in extended settings—weighted finite automata.
    • Potential tools: neural recognizers for protocol traces, differentiable parsers, runtime monitors, and learned controllers with explicit state-transition structure.
    • Dependencies: the strongest weighted-automaton and regular-language results require further assumptions, including additional layers, exact arithmetic or controlled precision, and in some cases expansive transitions with β>2\beta>2. These conditions may be unsuitable for numerically robust deployment.
  • Periodic and oscillatory control in robotics (robotics, autonomous systems, prosthetics). CKDA-like transitions could represent rotational phase and cyclic behaviors in locomotion, manipulation, gait generation, or repeated inspection trajectories.
    • Potential workflow: use a CKDA state as a learned phase memory, decode it into control targets, and combine it with a safety-constrained controller.
    • Dependencies: the paper demonstrates sequence prediction rather than closed-loop control. Stability under feedback, disturbances, actuator saturation, delay, and distribution shift must be demonstrated before deployment in safety-critical robotics.
  • Signal processing and forecasting of periodic physical systems (energy, manufacturing, transportation). The rotation-capable recurrence could model periodic demand, turbine vibrations, motor currents, grid-frequency fluctuations, traffic cycles, or seasonal environmental measurements.
    • Potential products: low-memory anomaly detectors, phase-aware forecasters, and streaming digital-twin components.
    • Dependencies: real systems often contain changing frequencies, harmonics, shocks, and trend components. CKDA may need multi-head or hybrid designs, explicit noise handling, and calibration; the paper does not establish advantages on real industrial datasets.
  • Memory-efficient audio and music generation (creative tools, speech, media). CKDA could support long-horizon continuation of rhythmically or harmonically structured signals, potentially reducing inference memory compared with attention-based generators.
    • Potential tools: loop extension, accompaniment generation, rhythm tracking, audio restoration, and real-time music effects.
    • Dependencies: the reported audio experiment uses a synthetic groove and open-loop waveform continuation. High-quality music generation requires modeling hierarchical structure, semantics, perceptual quality, and exposure to real audio distributions; nonlinear or attention-based models may remain necessary.
  • Theory-guided recurrent architectures with stronger state representations (academic research). CKDA provides a framework for studying how signed diagonal gates and Householder reflections create noncommuting, rotation-like dynamics within efficient real-valued RNNs.
    • Research directions: multi-plane rotations, more expressive low-rank transitions, learned orthogonal representations, improved initialization near known group constructions, and methods for correcting accumulated numerical errors.
    • Dependencies: the paper proves that one CKDA transition has at most one non-real conjugate eigenvalue pair and shows a one-layer obstruction for S_5 under stated assumptions. More expressive tasks may require multiple layers, more heads, auxiliary memory, or different transition families.
  • Policy and benchmarking standards for efficient sequence models (AI evaluation and public-sector procurement). The paper motivates evaluating recurrent models not only by perplexity but also by extrapolation, state-tracking, periodic continuation, numerical stability, and resource consumption.
    • Potential output: standardized benchmark suites covering group-word tasks, long-horizon periodic prediction, real streaming workloads, throughput, energy per token, and failure under state perturbations.
    • Dependencies: symbolic benchmarks can overstate practical capability, while language-model averages may conceal long-context failures. Evaluation should include real-world datasets, multiple precisions, adversarial perturbations, and comparisons with optimized attention–recurrent hybrids.
  • Safety-critical deployment in healthcare and finance (healthcare monitoring, financial time series) should be considered only as a longer-term possibility. CKDA’s streaming and periodic-state properties could eventually support continuous monitoring or low-latency forecasting, but the paper does not provide evidence for clinical diagnosis, financial prediction, calibration, fairness, or regulatory compliance.
    • Dependencies: extensive domain validation, uncertainty estimation, auditability, privacy protection, robustness testing, and human oversight would be mandatory. The current findings support architectural experimentation, not direct deployment in high-stakes decisions.

Glossary

  • Affine state update: A state-transition rule consisting of a linear transformation plus an additive term. “Linear recurrent neural networks process sequences through stacked layers with affine state updates.”
  • Algebraic multiplicity: The number of times an eigenvalue occurs as a root of a matrix’s characteristic polynomial. “at most one non-real conjugate eigenvalue pair on the unit circle, counted with algebraic multiplicity.”
  • Automaton emulation: Reproducing the state transitions and outputs of an automaton with another computational model. “selective SSM parameterizations for automaton emulation”
  • Bilinear interaction: An interaction that is linear in each of two arguments separately but may be nonlinear jointly. “Alternative transitions use bilinear interactions”
  • Chunk-wise implementation: A computation strategy that divides a sequence into blocks processed using specialized matrix operations. “efficient WY-based chunk-wise implementations.”
  • Commutator: The matrix measuring the failure of two matrices to commute, usually defined as AB−BAAB-BA. “Their commutator $=_1_2_1^{-1}_2^{-1}$ satisfies”
  • Complex-conjugate eigenvalue pair: A pair of non-real eigenvalues that are conjugates of one another and occur for real matrices. “a complex-conjugate pair can exist.”
  • Complex eigenvalue: An eigenvalue with a nonzero imaginary component, potentially representing rotational dynamics. “The resulting transition spectra show complex eigenvalues.”
  • Conjugacy: A relation between group elements or matrices formed by transforming one with an inverse and another element. “its conjugate by a transposition”
  • Continuous relaxation: A differentiable parameterization that allows optimization over a continuous range while representing a bounded or discrete target quantity. “our implementation uses βt=2σ(bt)\beta_t=2\sigma(b_t) and a signed gate based on rt,i=2σ(at,i)−1r_{t,i}=2\sigma(a_{t,i})-1”
  • D-dimensional orthogonal representation: A representation of abstract group elements as dd-dimensional orthogonal matrices that preserve inner products. “higher-dimensional orthogonal representations for permutation composition.”
  • Decoder: A function that maps a model’s hidden state to a task-relevant output or symbolic state. “a fixed decoder ff recovers its state from the model's hidden state hth_t”
  • Discriminant: A scalar expression that determines properties of a polynomial’s roots, such as whether a quadratic has real or complex roots. “its discriminant is negative”
  • Diagonal-plus-low-rank (DPLR): A matrix structure formed by adding a low-rank matrix to a diagonal matrix. “The planar construction shows how signed gating enables rotations within a rank-one diagonal-plus-low-rank (DPLR) transition.”
  • Delta-rule: A recurrent update using an identity-like transformation with a rank-one correction that selectively modifies the state. “non-diagonal transitions using the delta-rule introduce a rank-one correction”
  • Eigenvalue spectrum: The collection of eigenvalues of a matrix, including their multiplicities. “standard nonnegative KDA still has a real spectrum.”
  • Expansive transition: A state transition that can increase vector norms rather than preserve or contract them. “Both allow expansive transitions by default”
  • Faithful orthogonal representation: A group representation in which distinct group elements map to distinct orthogonal matrices. “A faithful orthogonal representation assigns each group element a distinct orthogonal matrix”
  • Finite reachability: The property that a recurrent system can visit only finitely many states under the considered updates. “Finite reachability (the set of RNN states is finite)”
  • Finite-state automaton: A computational model with finitely many internal states and transitions determined by input symbols. “deterministic automata, and weighted finite automata (WFAs).”
  • Fixed exact datatype: A numerical representation in which computations are performed with exact, rather than approximate, arithmetic. “These constructions use orthogonal transitions and a fixed exact datatype.”
  • Formal language: A set of strings defined by formal rules and recognized by computational models such as automata. “Finite monoids recognize exactly the regular languages”
  • Householder reflection: A reflection across a hyperplane, typically represented by a matrix of the form I−2vv⊤I-2vv^\top for a unit vector vv. “βi=0,1,2\beta_i=0,1,2 give identity, projection, and Householder reflection”
  • Identity-plus-rank-one matrix: A matrix formed by adding a rank-one update to the identity matrix. “each an identity-plus-rank-one matrix.”
  • Length extrapolation: The ability of a sequence model to generalize to sequences longer than those used during training. “yields the strongest length extrapolation”
  • Linear recurrent neural network (RNN): A recurrent neural network whose state update is linear or affine in the previous state. “Linear recurrent neural networks process sequences through stacked layers with affine state updates.”
  • Many-to-one decoder: A decoder that maps multiple distinct hidden states to the same output state. “a many-to-one decoder can map distinct hidden states to the same group element”
  • Non-expansiveness: The property that a transformation does not increase the norm or distance of its inputs. “It preserves KDA’s stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive”
  • Non-commutation: The failure of two mathematical operations or matrices to produce the same result when applied in either order. “Noncommutation alone, however, is not sufficient to produce a non-real spectrum.”
  • Non-real spectrum: A matrix spectrum containing eigenvalues that are not real numbers. “allowing the diagonal gate to change sign removes this restriction.”
  • Orthogonal matrix: A matrix whose transpose is its inverse, preserving lengths and angles. “Every orthogonal DPR1 matrix”
  • Periodic waveform continuation: Predicting the continuation of a repeating signal after an initial observed segment. “periodic waveform continuation”
  • Permutation group: A group whose elements are bijections of a set, with composition as the group operation. “Any finite group is isomorphic to a subgroup of a permutation group”
  • Polynomial precision: Numerical precision whose required number of bits grows polynomially with the problem size. “compute every WFA over Q\mathbb Q in polynomial precision”
  • Rank-one correction: A matrix update whose range has dimension one, allowing a structured modification of a base transformation. “introduce a rank-one correction that mixes information across state coordinates”
  • Regular language: A language recognizable by a finite automaton. “Three layers also recognize every regular language”
  • Representation discovery: The process by which a model learns internal structures corresponding to useful mathematical representations. “consistent with representation discovery being an optimization obstacle.”
  • Rotation: A norm-preserving transformation that changes the orientation of vectors, usually within a plane. “CKDA can realize 2D rotations”
  • Scalar gate: A gating mechanism that applies the same multiplicative value to every channel or coordinate. “Gated DeltaNet (GDN) uses a scalar gate”
  • Signed diagonal gate: A diagonal gating matrix whose entries may be positive or negative. “CKDA realizes planar rotations using gate entries of opposite sign”
  • State tracking: The task of maintaining and decoding a system’s evolving state as input-dependent transitions are composed over time. “We study this expressivity question through state tracking”
  • Transition matrix: A matrix that maps a recurrent state to its next state. “their linear updates with a low-rank correction constrain their expressivity.”
  • Unit circle: The set of complex numbers with magnitude one, often associated with stable oscillatory dynamics. “the complex pair moves outward toward the unit circle”
  • Weighted finite automaton (WFA): A finite-state automaton whose transitions and outputs carry numerical weights. “compute every WFA over Q\mathbb Q in polynomial precision”
  • Word problem: The problem of determining the group element represented by a sequence of group generators or inputs. “Three CKDA layers solve every finite group-word problem”
  • WY representation: A compact representation used to apply products of structured transformations efficiently, particularly Householder transformations. “both retain diagonal-plus-rank-one transitions and efficient WY-based chunk-wise implementations.”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 892 likes about this paper.