Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Abstract: Linear RNNs based on the delta-rule enable efficient sequence modeling, but their linear updates with a low-rank correction constrain their expressivity. Prior work has shown that composing two delta-rule transitions in a single recurrent update can model a 2D rotation, but this increases the rank and the cost of the updates compared to a single transition. We show that Kimi Delta Attention (KDA) can realize 2D rotations by combining a single delta-rule transformation with a second reflection supplied by its channel-wise gate. This requires extending the parameter ranges of KDA by combining two existing range extensions: allowing gates in and the delta-rule coefficient in . We call the resulting model Complex KDA (CKDA). It preserves KDA's stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive, while reaching the state-tracking expressivity of DeltaProduct. We characterize the expressivity of CKDA and prove that every orthogonal diagonal-plus-rank-one matrix is exactly a CKDA transition matrix. A single CKDA layer can track every finite group isomorphic to a subgroup of , and many state-tracking results use one fewer layer for CKDA compared to other diagonal-plus-rank-one Linear RNNs. Empirically, combining both extensions yields the strongest length extrapolation among tested KDA range settings on , , and periodic audio continuation. In language modeling, CKDA outperforms Transformers and other linear RNNs, obtains similar results to a KDA baseline, and shows promising scaling behavior. Our code is open source at https://github.com/OpenEuroLLM/ComplexKDA and our models are available at https://huggingface.co/collections/openeurollm/complexkda.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces a new type of neural-network memory called Complex KDA, or CKDA.
Neural networks that work with sequences—such as sentences, music, or time series—need to remember what happened earlier. CKDA is designed to do this efficiently, especially when sequences are very long.
The researchers’ main idea is to improve Kimi Delta Attention (KDA) by allowing parts of its memory update to use negative values. This small-looking change lets CKDA perform something very important: rotations.
A rotation is useful because it allows a model to keep track of repeating patterns, such as:
- counting around a circle,
- remembering the order of objects being moved,
- continuing a repeating musical rhythm,
- tracking complicated combinations of actions.
2. What questions did the researchers ask?
The paper focuses on several main questions:
- Can KDA remember more complicated patterns without becoming much slower?
- Can one KDA update perform a rotation, instead of needing two separate updates?
- Does allowing negative “gates” make the model more powerful?
- Can CKDA remember information over much longer sequences than ordinary KDA?
- Does this extra mathematical ability improve real tasks such as language modeling and audio continuation?
Here, a gate is like a control knob. It decides how strongly different parts of the model’s memory should be kept, changed, or reversed.
3. How did the researchers study this?
The researchers used both mathematics and computer experiments.
Mathematical analysis
They examined the rule CKDA uses to update its memory. In simplified form, the update is:
The delta update changes the memory using information from the current input. The channel-wise gate gives each memory channel its own control value.
A useful analogy is a group of arrows on a piece of paper:
- A normal update might stretch or shrink the arrows.
- A reflection flips them across a line.
- Two reflections at different angles create a rotation.
The researchers showed that CKDA can combine:
- a reflection controlled by the delta rule, and
- another reflection created by signed gates.
Together, these can produce rotations while keeping the update efficient.
They also studied the update matrices. A matrix is simply a table of numbers that describes how a model changes its memory. The researchers proved that CKDA can represent every orthogonal matrix of a certain useful form. Orthogonal means that the transformation preserves lengths, like a rotation or reflection.
State-tracking experiments
The researchers tested whether CKDA could remember the results of sequences of actions. For example, imagine a shell game with three objects. If objects are swapped many times, the model must remember where each object ended up.
They tested group-tracking problems involving:
- : arrangements of three objects,
- : arrangements of four objects,
- : a more difficult mathematical structure related to the symmetries of an icosahedron.
They trained the models on shorter sequences and then tested them on longer sequences. This is called length extrapolation: seeing whether a model can work beyond the sequence lengths it practiced.
Audio experiment
The researchers also gave the models part of a repeating musical waveform and asked them to continue it after the input stopped.
This tests whether the model can remember the phase of a repeating pattern—essentially, where it is in the cycle.
Language-modeling experiments
Finally, they trained large models with about 1.3 billion parameters on 100 billion tokens of educational web text. They compared CKDA with:
- Transformers,
- Mamba-style models,
- Gated DeltaNet,
- ordinary KDA,
- other recurrent neural networks.
They measured language-model quality, reasoning-task accuracy, and computation speed.
4. What did they find?
CKDA can perform rotations in one update
The most important theoretical result is that CKDA can create a 2D rotation using only one delta-rule update.
Ordinary KDA generally uses nonnegative gates. With only nonnegative gates, its transformations have limited behavior and cannot easily create true rotations.
CKDA allows:
- gates between and $1$,
- delta-rule strength between $0$ and $2$.
The negative gates can flip one part of the memory while leaving another part unchanged. Combined with another reflection, this creates a rotation.
This gives CKDA more expressive power without requiring two separate delta updates.
It can track important groups with one layer
The theory shows that a single CKDA layer can track:
- cyclic patterns such as counting around a loop,
- dihedral patterns involving rotations and reflections,
- the symmetry groups and ,
- and, with a slightly larger representation, .
In simpler terms, CKDA can remember complicated order-sensitive combinations of actions using fewer layers than several competing recurrent models.
However, the researchers also proved a limit: one CKDA layer cannot track under their stability conditions. This is important because it shows that CKDA is more powerful, but not unlimited.
Better long-sequence state tracking
On the and tests, CKDA was much better at remembering information over long sequences than the other tested KDA settings.
Using only one of the two improvements—negative gates or the larger delta range—was not enough. The strongest results came from using both together.
The model also learned the behavior predicted by the mathematics:
- delta strength close to ,
- gates close to either or ,
- complex-valued eigenvalues.
An eigenvalue is a mathematical way to describe how a transformation behaves. Complex eigenvalues often indicate that part of the system is rotating rather than simply growing or shrinking.
Strong periodic-audio continuation
CKDA was especially good at continuing the repeating audio waveform.
At a test length of 264 steps—well beyond the longest training length of 136 steps—CKDA achieved a signal-to-noise ratio of 38.1 dB. The causal Transformer achieved only 2.8 dB on this test.
This suggests that rotations help the model preserve repeating patterns over long periods.
The GRU still performed better on this particular audio task, so CKDA was not the best model in every situation.
Competitive language modeling
In language modeling, CKDA performed about as well as ordinary KDA and better than the Transformer and other recurrent models included in the comparison.
For example, on the reported language benchmarks:
| Model | WikiText perplexity | LAMBADA perplexity | Average reasoning accuracy |
|---|---|---|---|
| Transformer hybrid | 19.22 | 13.72 | 50.86% |
| Gated DeltaNet | 16.40 | 11.89 | 52.07% |
| KDA | 16.81 | 11.68 | 52.28% |
| CKDA | 15.78 | 10.08 | 54.06% |
Lower perplexity is better because it means the model is less surprised by the text. Higher reasoning accuracy is better.
CKDA also kept about 96–97% of standard KDA’s processing speed, so the additional expressive power did not cause a major slowdown.
5. Why is this important?
Many sequence models face a trade-off:
- Simple models are fast but may struggle with complicated patterns.
- More powerful models can remember more, but may require more computation or memory.
CKDA tries to improve this trade-off. It adds a small change—signed, channel-wise gates—but gains the ability to represent rotations and more complicated state changes.
This could be useful for:
- long-context LLMs,
- speech and music processing,
- periodic signals,
- tracking actions and object movements,
- efficient models for devices with limited memory or computing power.
The work also gives researchers a clearer understanding of why some recurrent models are more expressive than others. It shows that negative gates are not merely a minor numerical detail: they can change the geometry of the memory updates and allow the model to represent rotations.
Still, the results should be interpreted carefully. The strongest theoretical results rely on exact or carefully controlled mathematical settings, and some experiments use special initialization or simplified tasks. CKDA also does not solve every state-tracking problem, and its language-modeling improvement over KDA is modest.
Overall, the paper argues that Complex KDA is an efficient recurrent architecture that can remember certain long-term and repeating patterns better than standard KDA, while remaining fast enough for large-scale use.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Optimization versus representational capacity remains unresolved: CKDA can represent rotations and group actions theoretically, but standard training fails on without theory-specific initialization; the causes of this optimization barrier and methods for overcoming it are not systematically studied.
- The benefit of signed gates is not isolated in language modeling: CKDA performs similarly to bounded-gate KDA, but the experiments do not establish whether its additional rotational expressivity improves language modeling beyond changes in initialization, parameterization, optimization, or implementation.
- Scaling behavior is insufficiently characterized: Language-model results are reported at only approximately $1.3$B parameters and $100$B tokens, leaving open whether CKDA retains an advantage at substantially larger model sizes, longer training budgets, and different data mixtures.
- The comparison with Transformer and recurrent baselines is not fully controlled: The models may differ in training recipes, architectures, tokenization, normalization, parameter allocation, and implementation details, making it difficult to attribute performance differences specifically to CKDA’s transition structure.
- Long-context language modeling is not directly evaluated: The paper provides synthetic extrapolation and standard benchmark results, but does not measure perplexity, retrieval, or reasoning performance as a function of context length substantially beyond the training context.
- Numerical robustness of signed, near-orthogonal dynamics is unclear: Theoretical constructions use exact or fixed-precision arithmetic, whereas practical experiments use BF16; the accumulation of phase, sign, and amplitude errors over very long sequences is not quantitatively analyzed.
- Error-correction capability is unexplored: CKDA can preserve rotations, but the paper does not determine whether it can recover from noisy, corrupted, or ambiguous state updates, which is important for robust long-horizon tracking.
- The effect of gate magnitudes away from is underexplored: The theory emphasizes nearly sign-valued gates and , but it is unclear how performance changes across the continuous parameter ranges and , or whether intermediate values provide useful stability–adaptivity trade-offs.
- The interaction between non-expansiveness and practical forgetting is unresolved: Orthogonal transitions preserve information but do not inherently provide selective forgetting; the paper does not characterize how CKDA balances persistent memory with rapid state reset in realistic tasks.
- The claimed single-layer expressivity bounds are not known to be tight: The paper proves that one layer can track several groups and that one-layer tracking is impossible under specified assumptions, but the exact boundary for other finite groups, especially for , remains open.
- The impossibility result depends on restrictive assumptions: It assumes non-expansive transitions, finite reachability, and a finite number of independent heads; it remains unknown whether one-layer CKDA can track with expansive transitions, infinite or approximate reachability, shared cross-head state, or nonstandard decoders.
- Approximate state tracking is not theoretically characterized: The impossibility results concern exact decoding and finite reachability, whereas practical models operate with approximate representations; robustness thresholds and accuracy–length trade-offs are not derived.
- The role of multiple heads is incompletely understood: The paper rules out under independent-head assumptions but does not determine how shared parameters, cross-head mixing, head-specific dimensions, or nonlinear inter-head decoders alter expressivity.
- The minimal dimension and layer count are unknown: The constructions establish sufficient dimensions and depths for several groups and automata, but do not generally provide minimal state dimension, number of heads, or number of CKDA layers.
- The relationship between CKDA products and minimal factorization complexity is incomplete: Upper bounds on the number of CKDA factors are given, but the exact minimum number of factors required for important classes of orthogonal or non-expansive matrices is not established beyond the stated sharp orthogonal bound.
- The scope of the one-complex-pair limitation is unclear for learned products: Each CKDA transition has at most one non-real conjugate eigenvalue pair, while products can be universal, but the paper does not characterize how many steps, layers, or factors are needed to generate specified multi-frequency dynamics.
- Generalization beyond finite groups is limited: The empirical state-tracking evaluation focuses mainly on , , and ; performance on non-group finite automata, semigroups, modular arithmetic with noisy inputs, hierarchical counters, and compositional algorithmic tasks remains untested.
- Theoretical constructions for regular languages and weighted finite automata are not empirically validated: The paper proves multilayer results, including constructions requiring , but does not demonstrate these capabilities on representative regular-language or WFA computation benchmarks.
- The practical consequences of allowing are unresolved: Expansive transitions are required for some theoretical constructions, yet their numerical stability, training dynamics, precision requirements, and effect on long-context language modeling are not experimentally assessed.
- The audio experiment has limited breadth: It uses a single synthetic periodic groove and a single-forward-pass continuation setup; it does not establish whether CKDA benefits real music, speech, multiscale periodic signals, irregular frequencies, changing tempo, or noisy observations.
- The advantage over nonlinear recurrent models is not established: The GRU outperforms CKDA on the reported audio task, but the paper does not investigate whether hybrid CKDA–nonlinear architectures can combine rotational extrapolation with nonlinear phase inference.
- The effect of input and output nonlinearities is insufficiently separated: State-tracking experiments remove the inherited SiLU activation, whereas language-model experiments ablate it only partially; the role of key/query nonlinearities in enabling or suppressing rotational dynamics remains unclear.
- Initialization sensitivity is not systematically measured: The paper uses spread signed-gate initialization and specialized initialization for , but does not report success rates across seeds, initialization scales, gate-sign proportions, or training schedules.
- Throughput and memory claims lack broader systems evaluation: The reported $96$– of KDA throughput is measured for a specific H100 configuration and sequence shape; scaling across GPUs, sequence lengths, batch sizes, head dimensions, backward passes, memory usage, and distributed training is not established.
- The cost of signed-gate handling during inference is not fully quantified: The paper discusses gauge transformations and fused kernels, but does not provide end-to-end latency, energy, memory-bandwidth, or deployment costs relative to KDA and DeltaProduct.
- The stability analysis is primarily spectral and structural: Non-expansive transitions are shown to preserve norm bounds, but the paper does not analyze transient amplification in non-normal products, gradient stability, sensitivity to perturbations, or stability of the complete nonlinear input-dependent network.
- The decoder’s role may overstate practical tracking ability: Several theoretical results permit sufficiently expressive feed-forward or many-to-one decoders; the complexity, parameter count, training feasibility, and generalization of these decoders are not bounded.
- The relationship between learned CKDA transitions and the prescribed group representations is only qualitatively examined: The interpretability analysis shows complex spectra and near-reflection parameters for , but does not quantify representation alignment, transition error, subgroup closure, or how consistently the learned mechanism matches the theoretical construction.
- No systematic robustness evaluation is provided: Effects of token noise, missing inputs, adversarial perturbations, quantization, reduced precision, sequence truncation, and distribution shift on state tracking and language modeling remain unknown.
- The usefulness of CKDA for tasks requiring non-periodic and content-dependent memory is uncertain: The strongest evidence concerns rotations and periodic continuation, so it remains unclear whether signed rotational dynamics improve retrieval, induction, copying, hierarchical reasoning, or other non-periodic sequence behaviors.
Practical Applications
Immediate Applications
- Deploy CKDA as a long-context language-model backbone (software, NLP, enterprise AI). The released implementation and models can be used to build recurrent LLMs for document processing, retrieval-augmented generation, code completion, and streaming text generation. CKDA is particularly relevant where linear-time sequence processing and fixed-size recurrent state are preferable to quadratic self-attention.
- Practical workflow: replace or augment KDA/DeltaNet blocks in an existing decoder-only LLM, fine-tune on domain data, and benchmark perplexity, throughput, memory use, and long-context accuracy.
- Evidence from the paper: CKDA retains approximately
96–97%of standard KDA throughput and performs comparably to KDA while outperforming the listed Transformer and linear-RNN baselines in the reported 1.3B-parameter experiment. - Dependencies: results depend on model scale, training data, initialization, hardware kernels, numerical precision, and whether signed gates are successfully optimized. The reported results do not establish superiority across all tasks or scales.
- Efficient streaming inference for long sequences (edge computing, communications, monitoring, embedded AI). CKDA’s fixed-size recurrent state and linear sequence scaling can support online processing of sensor streams, logs, transcripts, and user interactions without storing the entire history.
- Potential products: streaming speech or text assistants, online log summarizers, anomaly-monitoring agents, and low-memory document or message processors.
- Why it is actionable: the recurrence is compatible with existing KDA-style kernels and does not require an auxiliary recurrent state or explicit phase representation.
- Dependencies: recurrent-state compression may lose information on tasks requiring arbitrary retrieval; production systems should compare CKDA against attention or hybrid architectures for factual recall and error accumulation.
- Periodic signal continuation and phase tracking (audio, industrial monitoring, robotics, control). CKDA can be used for extrapolating periodic or quasi-periodic signals after an observed prefix, such as audio waveforms, vibration signals, rotating-machine telemetry, or cyclic actuator trajectories.
- Potential workflow: infer the phase and latent state from a short cue, then generate or forecast future values in a single forward pass.
- Evidence from the paper: CKDA achieved
38.1 dBSNR at a continuation length beyond the training horizon on the synthetic periodic-groove task, whereas the causal Transformer degraded substantially. - Dependencies: the experiment is synthetic and periodic; performance on noisy, drifting, multi-frequency, or nonstationary real-world signals remains unverified. A GRU was more accurate on the reported task, so CKDA should be treated as an efficient alternative rather than a universally superior forecaster.
- Long-horizon symbolic and permutation-state tracking (education technology, software testing, algorithmic reasoning). CKDA can serve as a compact recurrent model for tasks involving parity, modular counting, cyclic transformations, and permutation composition.
- Potential tools: sequence-learning benchmarks, differentiable finite-state machines, automated reasoning modules, and models that track object swaps or action sequences.
- Evidence from the paper: a single CKDA layer tracks finite cyclic and dihedral groups and successfully extrapolates on the reported
S_3andS_4tasks. - Dependencies: exact or sufficiently stable numerical representations, a suitable decoder, and training procedures that discover the rotational mechanism are important. These benchmark capabilities do not automatically imply robust general reasoning in unconstrained environments.
- Use as an experimental drop-in for existing KDA systems (academic and industrial ML engineering). The open-source implementation can support controlled ablations of gate ranges and delta-rule coefficients.
- Actionable comparison: evaluate standard KDA, bounded positive-gate KDA, signed-gate KDA, and CKDA on the same data, measuring extrapolation, throughput, memory, numerical stability, and optimization behavior.
- Dependencies: signed-gate implementations require correct sign handling in forward and backward passes; kernel support and hardware compatibility may affect the claimed throughput.
- Compact recurrent modeling for on-device applications (mobile, IoT, robotics). Because CKDA uses diagonal-plus-rank-one transitions and non-expansive updates, it is a candidate for memory-constrained inference in devices processing continuous input.
- Potential applications: wearable-sensor interpretation, robot activity recognition, predictive maintenance, and local voice or gesture processing.
- Dependencies: the paper evaluates GPU kernels rather than complete low-power deployments. Quantization, power consumption, latency under batch size one, and robustness to sensor noise require separate validation.
Long-Term Applications
- Long-context foundation models with recurrent or hybrid architectures (NLP, multimodal AI). CKDA could become a recurrent alternative to attention in models that need to process very long documents, code repositories, genomic sequences, or multimodal event streams.
- Potential architecture: use CKDA for the bulk of sequence processing and reserve attention or external memory for retrieval-heavy operations.
- Expected benefit: learned rotational dynamics may preserve phase, order, and periodic structure over longer horizons while keeping linear inference cost.
- Dependencies: scaling behavior beyond the reported 1.3B-parameter experiment must be established. Important unresolved issues include optimization at large scale, recurrent-state corruption, token-level retrieval, parallel training efficiency, and performance on real long-context benchmarks.
- Robust finite-state and automaton emulation (formal methods, verification, program synthesis). The paper’s group-tracking results suggest using CKDA as a differentiable implementation of finite-state machines, regular languages, and—in extended settings—weighted finite automata.
- Potential tools: neural recognizers for protocol traces, differentiable parsers, runtime monitors, and learned controllers with explicit state-transition structure.
- Dependencies: the strongest weighted-automaton and regular-language results require further assumptions, including additional layers, exact arithmetic or controlled precision, and in some cases expansive transitions with . These conditions may be unsuitable for numerically robust deployment.
- Periodic and oscillatory control in robotics (robotics, autonomous systems, prosthetics). CKDA-like transitions could represent rotational phase and cyclic behaviors in locomotion, manipulation, gait generation, or repeated inspection trajectories.
- Potential workflow: use a CKDA state as a learned phase memory, decode it into control targets, and combine it with a safety-constrained controller.
- Dependencies: the paper demonstrates sequence prediction rather than closed-loop control. Stability under feedback, disturbances, actuator saturation, delay, and distribution shift must be demonstrated before deployment in safety-critical robotics.
- Signal processing and forecasting of periodic physical systems (energy, manufacturing, transportation). The rotation-capable recurrence could model periodic demand, turbine vibrations, motor currents, grid-frequency fluctuations, traffic cycles, or seasonal environmental measurements.
- Potential products: low-memory anomaly detectors, phase-aware forecasters, and streaming digital-twin components.
- Dependencies: real systems often contain changing frequencies, harmonics, shocks, and trend components. CKDA may need multi-head or hybrid designs, explicit noise handling, and calibration; the paper does not establish advantages on real industrial datasets.
- Memory-efficient audio and music generation (creative tools, speech, media). CKDA could support long-horizon continuation of rhythmically or harmonically structured signals, potentially reducing inference memory compared with attention-based generators.
- Potential tools: loop extension, accompaniment generation, rhythm tracking, audio restoration, and real-time music effects.
- Dependencies: the reported audio experiment uses a synthetic groove and open-loop waveform continuation. High-quality music generation requires modeling hierarchical structure, semantics, perceptual quality, and exposure to real audio distributions; nonlinear or attention-based models may remain necessary.
- Theory-guided recurrent architectures with stronger state representations (academic research). CKDA provides a framework for studying how signed diagonal gates and Householder reflections create noncommuting, rotation-like dynamics within efficient real-valued RNNs.
- Research directions: multi-plane rotations, more expressive low-rank transitions, learned orthogonal representations, improved initialization near known group constructions, and methods for correcting accumulated numerical errors.
- Dependencies: the paper proves that one CKDA transition has at most one non-real conjugate eigenvalue pair and shows a one-layer obstruction for
S_5under stated assumptions. More expressive tasks may require multiple layers, more heads, auxiliary memory, or different transition families.
- Policy and benchmarking standards for efficient sequence models (AI evaluation and public-sector procurement). The paper motivates evaluating recurrent models not only by perplexity but also by extrapolation, state-tracking, periodic continuation, numerical stability, and resource consumption.
- Potential output: standardized benchmark suites covering group-word tasks, long-horizon periodic prediction, real streaming workloads, throughput, energy per token, and failure under state perturbations.
- Dependencies: symbolic benchmarks can overstate practical capability, while language-model averages may conceal long-context failures. Evaluation should include real-world datasets, multiple precisions, adversarial perturbations, and comparisons with optimized attention–recurrent hybrids.
- Safety-critical deployment in healthcare and finance (healthcare monitoring, financial time series) should be considered only as a longer-term possibility. CKDA’s streaming and periodic-state properties could eventually support continuous monitoring or low-latency forecasting, but the paper does not provide evidence for clinical diagnosis, financial prediction, calibration, fairness, or regulatory compliance.
- Dependencies: extensive domain validation, uncertainty estimation, auditability, privacy protection, robustness testing, and human oversight would be mandatory. The current findings support architectural experimentation, not direct deployment in high-stakes decisions.
Glossary
- Affine state update: A state-transition rule consisting of a linear transformation plus an additive term. “Linear recurrent neural networks process sequences through stacked layers with affine state updates.”
- Algebraic multiplicity: The number of times an eigenvalue occurs as a root of a matrix’s characteristic polynomial. “at most one non-real conjugate eigenvalue pair on the unit circle, counted with algebraic multiplicity.”
- Automaton emulation: Reproducing the state transitions and outputs of an automaton with another computational model. “selective SSM parameterizations for automaton emulation”
- Bilinear interaction: An interaction that is linear in each of two arguments separately but may be nonlinear jointly. “Alternative transitions use bilinear interactions”
- Chunk-wise implementation: A computation strategy that divides a sequence into blocks processed using specialized matrix operations. “efficient WY-based chunk-wise implementations.”
- Commutator: The matrix measuring the failure of two matrices to commute, usually defined as . “Their commutator $=_1_2_1^{-1}_2^{-1}$ satisfies”
- Complex-conjugate eigenvalue pair: A pair of non-real eigenvalues that are conjugates of one another and occur for real matrices. “a complex-conjugate pair can exist.”
- Complex eigenvalue: An eigenvalue with a nonzero imaginary component, potentially representing rotational dynamics. “The resulting transition spectra show complex eigenvalues.”
- Conjugacy: A relation between group elements or matrices formed by transforming one with an inverse and another element. “its conjugate by a transposition”
- Continuous relaxation: A differentiable parameterization that allows optimization over a continuous range while representing a bounded or discrete target quantity. “our implementation uses and a signed gate based on ”
- D-dimensional orthogonal representation: A representation of abstract group elements as -dimensional orthogonal matrices that preserve inner products. “higher-dimensional orthogonal representations for permutation composition.”
- Decoder: A function that maps a model’s hidden state to a task-relevant output or symbolic state. “a fixed decoder recovers its state from the model's hidden state ”
- Discriminant: A scalar expression that determines properties of a polynomial’s roots, such as whether a quadratic has real or complex roots. “its discriminant is negative”
- Diagonal-plus-low-rank (DPLR): A matrix structure formed by adding a low-rank matrix to a diagonal matrix. “The planar construction shows how signed gating enables rotations within a rank-one diagonal-plus-low-rank (DPLR) transition.”
- Delta-rule: A recurrent update using an identity-like transformation with a rank-one correction that selectively modifies the state. “non-diagonal transitions using the delta-rule introduce a rank-one correction”
- Eigenvalue spectrum: The collection of eigenvalues of a matrix, including their multiplicities. “standard nonnegative KDA still has a real spectrum.”
- Expansive transition: A state transition that can increase vector norms rather than preserve or contract them. “Both allow expansive transitions by default”
- Faithful orthogonal representation: A group representation in which distinct group elements map to distinct orthogonal matrices. “A faithful orthogonal representation assigns each group element a distinct orthogonal matrix”
- Finite reachability: The property that a recurrent system can visit only finitely many states under the considered updates. “Finite reachability (the set of RNN states is finite)”
- Finite-state automaton: A computational model with finitely many internal states and transitions determined by input symbols. “deterministic automata, and weighted finite automata (WFAs).”
- Fixed exact datatype: A numerical representation in which computations are performed with exact, rather than approximate, arithmetic. “These constructions use orthogonal transitions and a fixed exact datatype.”
- Formal language: A set of strings defined by formal rules and recognized by computational models such as automata. “Finite monoids recognize exactly the regular languages”
- Householder reflection: A reflection across a hyperplane, typically represented by a matrix of the form for a unit vector . “ give identity, projection, and Householder reflection”
- Identity-plus-rank-one matrix: A matrix formed by adding a rank-one update to the identity matrix. “each an identity-plus-rank-one matrix.”
- Length extrapolation: The ability of a sequence model to generalize to sequences longer than those used during training. “yields the strongest length extrapolation”
- Linear recurrent neural network (RNN): A recurrent neural network whose state update is linear or affine in the previous state. “Linear recurrent neural networks process sequences through stacked layers with affine state updates.”
- Many-to-one decoder: A decoder that maps multiple distinct hidden states to the same output state. “a many-to-one decoder can map distinct hidden states to the same group element”
- Non-expansiveness: The property that a transformation does not increase the norm or distance of its inputs. “It preserves KDA’s stability and efficiency, with transitions that remain diagonal-plus-rank-one and non-expansive”
- Non-commutation: The failure of two mathematical operations or matrices to produce the same result when applied in either order. “Noncommutation alone, however, is not sufficient to produce a non-real spectrum.”
- Non-real spectrum: A matrix spectrum containing eigenvalues that are not real numbers. “allowing the diagonal gate to change sign removes this restriction.”
- Orthogonal matrix: A matrix whose transpose is its inverse, preserving lengths and angles. “Every orthogonal DPR1 matrix”
- Periodic waveform continuation: Predicting the continuation of a repeating signal after an initial observed segment. “periodic waveform continuation”
- Permutation group: A group whose elements are bijections of a set, with composition as the group operation. “Any finite group is isomorphic to a subgroup of a permutation group”
- Polynomial precision: Numerical precision whose required number of bits grows polynomially with the problem size. “compute every WFA over in polynomial precision”
- Rank-one correction: A matrix update whose range has dimension one, allowing a structured modification of a base transformation. “introduce a rank-one correction that mixes information across state coordinates”
- Regular language: A language recognizable by a finite automaton. “Three layers also recognize every regular language”
- Representation discovery: The process by which a model learns internal structures corresponding to useful mathematical representations. “consistent with representation discovery being an optimization obstacle.”
- Rotation: A norm-preserving transformation that changes the orientation of vectors, usually within a plane. “CKDA can realize 2D rotations”
- Scalar gate: A gating mechanism that applies the same multiplicative value to every channel or coordinate. “Gated DeltaNet (GDN) uses a scalar gate”
- Signed diagonal gate: A diagonal gating matrix whose entries may be positive or negative. “CKDA realizes planar rotations using gate entries of opposite sign”
- State tracking: The task of maintaining and decoding a system’s evolving state as input-dependent transitions are composed over time. “We study this expressivity question through state tracking”
- Transition matrix: A matrix that maps a recurrent state to its next state. “their linear updates with a low-rank correction constrain their expressivity.”
- Unit circle: The set of complex numbers with magnitude one, often associated with stable oscillatory dynamics. “the complex pair moves outward toward the unit circle”
- Weighted finite automaton (WFA): A finite-state automaton whose transitions and outputs carry numerical weights. “compute every WFA over in polynomial precision”
- Word problem: The problem of determining the group element represented by a sequence of group generators or inputs. “Three CKDA layers solve every finite group-word problem”
- WY representation: A compact representation used to apply products of structured transformations efficiently, particularly Householder transformations. “both retain diagonal-plus-rank-one transitions and efficient WY-based chunk-wise implementations.”

