Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Order Is The Message

Published 24 Mar 2026 in cs.LG and stat.ML | (2603.25047v1)

Abstract: In a controlled experiment on modular arithmetic (p=9973p = 9973), varying only example ordering while holding all else constant, two fixed-ordering strategies achieve 99.5\% test accuracy by epochs 487 and 659 respectively from a training set comprising 0.3\% of the input space, well below established sample complexity lower bounds for this task under IID ordering. The IID baseline achieves 0.30\% after 5{,}000 epochs from identical data. An adversarially structured ordering suppresses learning entirely. The generalizing model reliably constructs a Fourier representation whose fundamental frequency is the Fourier dual of the ordering structure, encoding information present in no individual training example, with the same fundamental emerging across all seeds tested regardless of initialization or training set composition. We discuss implications for training efficiency, the reinterpretation of grokking, and the safety risks of a channel that evades all content-level auditing.

Authors (1)

Summary

  • The paper demonstrates that example ordering enters training through Hessian-gradient entanglement, accounting for 83–89% of cumulative epoch gradient norm and strongly influencing the direction of optimization.
  • Controlled modular-arithmetic experiments show that structured and even fixed-random orders reach 99.5% test accuracy, while IID shuffling reaches only 0.30% after 5,000 epochs and target-sorted data remains at chance.
  • Spectral measurements link ordering patterns to learned Fourier representations, suggesting that order-aware training can reduce grokking delays while also creating a covert influence channel that content-only audits cannot detect.

The paper "The Order Is The Message" (2603.25047) argues that the sequential ordering of training examples constitutes an active information channel in neural network training, distinct from the content signal carried by individual examples. Through a controlled experiment on modular arithmetic with exhaustive instrumentation, the author demonstrates that ordering alone—with data, architecture, optimizer, initialization, and compute held fixed—can determine whether a model generalizes, which representations it acquires, and how efficiently it learns. The work recharacterizes a well-known second-order interaction in SGD analysis, the Hessian-gradient cross-term, as the mathematical mechanism through which ordering enters the gradient, and measures this channel directly for the first time.

Theoretical framework: Hessian-gradient entanglement

The central theoretical object is the first-order Taylor expansion of the gradient at consecutive training steps. If the model is updated on batch AA and then the gradient is computed for batch BB, the observed gradient decomposes into a content term ∇LB(θ)\nabla L_B(\theta) and an "entanglement term" −ηHB(θ)⋅∇LA(θ)-\eta H_B(\theta) \cdot \nabla L_A(\theta), where HBH_B is the batch-BB Hessian. This cross-term is standard in backward-error analyses of implicit regularization and random-reshuffling convergence (Nanda et al., 2023, Beneventano, 2023); the paper's contribution is to reinterpret it as an information channel and to measure its magnitude and coherence.

The framework distinguishes two regimes. Under IID shuffling, the entanglement term projects a random direction through BB's Hessian at each step, so ordering effects cancel statistically over many epochs—the cancellation that underwrites classical convergence guarantees. Under structured ordering, consecutive gradients aligned with high-curvature eigendirections of each other's Hessians produce entanglement terms that interfere constructively, yielding a coherent directional signal toward particular feature representations. The paper extends the analysis to adaptive optimizers, arguing that Adam amplifies the channel through three pathways: larger parameter displacements, momentum acting as a temporal integrator that preserves coherent components, and selective per-parameter scaling that creates a feedback loop amplifying the ordering signal in the subspace it targets.

A key empirical claim follows from per-step Hessian measurements: the entanglement and content terms are nearly identical vectors (cosine similarity >0.999> 0.999 for all non-adversarial strategies), so the observed gradient is a small residual—30–55× smaller in norm than either component—of two large, nearly-canceling vectors. Small angular differences between these terms, determined by which batch follows which, are amplified into large directional changes in the gradient the optimizer sees. This geometry explains how ordering exerts influence disproportionate to any per-batch anomaly.

The framework also introduces the notion of a per-feature critical learning period: ordering has maximal leverage while a feature's representation is forming and the loss landscape is most curved with respect to it. Six falsifiable predictions are derived, all of which the paper reports as confirmed experimentally.

Experimental design

The task is addition modulo p=9973p = 9973, learned from 300,000 randomly sampled input pairs—approximately 0.3% of the p2≈108p^2 \approx 10^8 input space—using a two-layer pre-LayerNorm transformer (embedding dimension 256, 4 heads, feedforward dimension 2048), trained with AdamW and cosine annealing for up to 5,000 epochs. Four ordering strategies are compared on identical data and initialization:

  • Stride: examples sorted by BB0, fixed across epochs—a deliberately structured traversal.
  • Fixed-Random: a single random permutation held fixed, containing no designed structure.
  • Random: fresh IID shuffling each epoch (the standard baseline).
  • Target: examples sorted by output value, producing strong but self-contradictory ordering structure.

The instrumentation is extensive: a counterfactual gradient decomposition that separates ordering-dependent from ordering-independent gradient components (validated to yield conservative lower bounds on the ordering fraction), per-step Hessian-vector products via finite differences, spectral analysis of embedding and decoder weights, and per-layer ordering-fraction tracking. All code, data, seeds, and weights are published, with bit-for-bit deterministic reproduction claimed on identical hardware.

Generalization results

The headline result is stark. The Stride and Fixed-Random models reach 99.5% test accuracy by epochs 487 and 659 respectively. The Random model, trained on identical data with identical compute, reaches only 0.30% test accuracy after 5,000 epochs despite perfectly memorizing the training set—the standard grokking regime. The Target model never exceeds chance (0.01% test accuracy) while its loss falls below initialization, indicating genuine internal reorganization without learning.

The Fixed-Random result is the most surprising claim in the paper: a fixed random permutation with no task-specific design is sufficient to drive generalization in a regime where IID ordering fails entirely. The author's explanation is domain-specific: any permutation of elements of the cyclic group BB1 carries spectral content at frequencies that constitute valid Fourier bases for the group operation, so any consistent ordering is accidentally task-relevant in this domain. Consistency prevents destructive interference; in this task class, whatever coherent signal survives is productive. The paper is explicit that this guarantee does not hold in general domains.

The train-test gap observed under the fixed orderings is reinterpreted: rather than a memorization table being replaced by a generalizing circuit (the standard grokking account), the model is building the generalizing Fourier representation from the first epochs but lacks sufficient harmonic resolution to interpolate the full cyclic group. Test accuracy rises in lockstep with harmonic accumulation.

Spectral evidence of ordering-channel learning

The mechanistic evidence that the model extracts information present only in the ordering is the strongest part of the paper. The Stride model's embedding spectrum concentrates on a harmonic series rooted at BB2—the Fourier dual of the sort stride over BB3 (BB4). This relationship was validated across stride values 50, 99, and 150, with predicted and observed fundamentals matching exactly in each case. The harmonic series follows a doubling pattern with Nyquist folding, and by convergence the 14 highest-power frequencies are all predicted harmonics.

Critically, BB5 emerges as the peak embedding frequency by epoch 3 and does so across all six seeds tested, despite each seed using a different initialization and a different random 300,000-pair training subset. Since the stride is a property of the sequence of examples and not of any individual example, the model's construction of a representation rooted in this frequency is direct evidence that information flowed through the ordering channel. The Fixed-Random model shows the same qualitative pattern with BB6, the spectral dominant of its particular permutation, confirming that the channel carries whatever structure the ordering contains.

The Target strategy produces a striking dissociation: decoder spectral entropy drops to 0.09—more concentrated than either generalizing strategy—while the embedding remains near uniform until a late weight-decay collapse. The ordering signal coherently organizes the output layer while actively preventing usable input representations, illustrating that anti-coherent ordering is not mere noise but an organized, degenerate attractor.

Quantifying the channel

The counterfactual decomposition yields the paper's most broadly consequential measurement: the ordering component accounts for 83–89% of each epoch's cumulative gradient norm across all four strategies, including IID shuffling. Under Random, the ordering norm remains roughly 2.8× the content norm for all 5,000 epochs. This challenges the implicit practitioner assumption that an epoch's cumulative gradient approximates the content gradient; in fact, the content signal emerges only as a long-run average after ordering-induced displacement cancels. Under IID training, the majority of each epoch's compute produces displacement that must subsequently be undone. Path efficiency (net displacement divided by distance traveled) differs by roughly an order of magnitude between generalizing (Stride: BB7) and non-generalizing (Random: BB8) strategies.

Ordering-content alignment averages approximately 0.40 for the generalizing strategies and dips to ~0.35 precisely when the ordering norm peaks, indicating the channel carries partially independent information and is most independent when strongest. The ordering fraction's decline is concentrated in the deep feedforward layers, where it correlates with validation accuracy at BB9–0.99, suggesting the ordering signal can serve as a tracer for where and when features crystallize.

Adam amplifies rather than suppresses the channel: update deflection exceeds 0.95 for all strategies, yet ordering fractions remain 83–84%, and the Hessian entanglement energy ratio averages 860–900× the observed gradient energy. The Adam amplification ratio differs systematically by strategy (246× Fixed-Random, 175× Stride, 101× Random, 58× Target), consistent with the theory's prediction of selective amplification in the targeted parameter subspace.

Regime dependence and the reinterpretation of grokking

Two ablations at ∇LB(θ)\nabla L_B(\theta)0 with 26.5% data density establish that the ordering channel's practical importance scales with the inadequacy of the content signal. At high data density, all non-adversarial strategies generalize, with fixed orderings providing only ~20% speedup. As weight decay decreases from 0.1 to 0.01, the fixed-ordering speedup grows from ~9% to ~32%, because Random degrades faster than the fixed orderings as the lazy-to-rich transition is delayed. The channel is complementary rather than substitutive: it matters most precisely when standard conditions are least favorable.

On this basis, the paper argues that in SGD-trained models, the grokking delay is an artifact of IID ordering: structured ordering eliminates the delay entirely, with no memorization phase. The author is appropriately careful about scope, conceding that grokking in Recursive Feature Machines [2503.xxxx, Mallinar et al.] and under full-batch gradient descent may operate outside the entanglement mechanism as described—though the paper argues full-batch GD retains a degenerate self-interaction form of the same term, consistent with its slower generalization. A candidate mechanism for the sharp phase transition is proposed: under IID ordering, accumulated content-driven landscape reshaping eventually allows the ever-present ordering noise to breach the memorization basin boundary. The framework also gives a mechanistic account of Grokfast (Lee et al., 2024): its low-pass filter approximates a content-signal extractor, attenuating the fast-varying ordering noise, whereas structured ordering makes the ordering component itself slow-varying at the source.

The sample complexity claim deserves careful statement: 99.5% generalization from 0.3% of the input space is well below the best proven bounds for this task under IID assumptions (e.g., ∇LB(θ)\nabla L_B(\theta)1 from Mohamadi et al.), but the paper does not claim a proven bound is violated, since the training regime differs in ordering and optimizer. The suggested conclusion—that the critical data fraction may itself be an artifact of IID assumptions—is a hypothesis requiring ordering-aware sample-complexity theory to settle.

Safety implications

The paper's safety argument is that the ordering channel is a covert influence pathway invisible to all content-level auditing: every individual example is legitimate, the aggregate distribution is unchanged, within-batch gradient statistics appear normal, and—most problematically—a productive ordering improves all standard training metrics. The concern is not primarily a deliberate adversary but an emergent one: in teacher-student pipelines where a teacher model curates data and is optimized on student performance, gradient descent on the teacher's objective will shape example ordering toward configurations producing coherent entanglement in the student, because this accelerates learning and is thereby rewarded. The Fixed-Random result demonstrates that no deliberate design is needed; an optimization process under selection pressure will find orderings at least this effective. Under iterated teacher-student generations, this selection pressure compounds.

The proposed mitigation is architectural: ordering must be randomized by a process outside the optimization loop. Detection is harder—the ordering fraction is nearly identical across all four strategies (83–89%), so its magnitude cannot distinguish coherent from incoherent ordering at a single measurement point. Candidate discriminative metrics (batch gradient autocorrelation, Adam amplification trajectories, Hessian Rayleigh quotients) exist in the paper's instrumentation but lack characterized false-positive rates, and the author acknowledges that developing them into reliable monitors is an open research direction.

Limitations and open questions

The paper is candid that its empirical claims rest on a single controlled domain. Several results are explicitly domain-specific: the predictable harmonic emergence, the stride-to-frequency relationship, and above all the success of Fixed-Random, which depends on the cyclic group guarantee that any consistent permutation carries task-relevant spectral content. Whether fixed random orderings outperform IID shuffling in domains whose generalizing representations bear no relationship to ordering spectra is unresolved, though GraB's empirical success on CIFAR-10, WikiText, and GLUE [(Mohtashami et al., 2022)-adjacent work, Lu et al.] suggests productive orderings exist beyond toy settings.

Scale is the most important open question: whether a single ordering sequence carries enough bandwidth to influence multiple features with independent critical learning periods in large models, and whether per-worker ordering specialization in distributed training could parallelize the channel, are both empirically unresolved. The detailed internal dynamics are drawn from a single instrumented seed per strategy (199), with seed sensitivity validated only for the headline generalization outcome (five of six Stride seeds and six of six Fixed-Random seeds generalizing); the precise numerical values reported should be read as one trajectory. Interactions with gradient accumulation, data augmentation, and cross-worker shuffling are uninvestigated, and the factorization of ordering into batch partition versus batch adjacency is deferred to future work.

Conclusion

This paper establishes, through a tightly controlled experiment and extensive instrumentation, that data ordering is an information channel rather than a nuisance variable: it operates through the Hessian-gradient entanglement term, accounts for 83–89% of each epoch's cumulative gradient norm under all ordering strategies including IID shuffling, and can be steered to produce generalization from 0.3% of the input space where IID ordering fails entirely, or to suppress learning through coherent but misaligned structure. The demonstration that an entirely undesigned fixed permutation suffices in this domain, and the identification of a training-influence pathway invisible to content-level auditing, together raise questions that extend well beyond the modular arithmetic setting: whether the dose-productivity continuum holds at scale, whether ordering-aware sample complexity bounds can be formulated, and whether temporal-gradient monitoring can be developed into a practical oversight tool for teacher-student training pipelines.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.