Papers
Topics
Authors
Recent
Search
2000 character limit reached

ByteGen: Byte-Level LOB Modeling

Updated 7 July 2026
  • ByteGen is a tokenizer-free generative model that represents limit order book events directly in byte space using a fixed 32-byte packed format.
  • It employs a hybrid Mamba–Transformer H-Net with dynamic chunking to discover latent segmentation without relying on predefined token boundaries.
  • Trained on 34.2 million CME Bitcoin futures events, ByteGen reproduces key market stylized facts, such as realistic price distributions and heavy-tailed returns.

Searching arXiv for ByteGen and closely related architecture papers to ground the article with current references. ByteGen is a tokenizer-free generative model for high-frequency limit order book event streams that operates directly in byte space rather than over engineered features or discrete symbolic tokens. It is introduced as the first end-to-end byte-level framework for LOB modeling and formulates market generation as autoregressive next-byte prediction over a fixed 32-byte packed representation of each event. The model combines a lossless 32-byte event encoding, a 256-symbol byte vocabulary, and a hybrid Mamba–Transformer H-Net with dynamic chunking that learns latent segmentation structure without predefined token boundaries. Trained on 34.2 million CME Bitcoin futures Level-3 events, ByteGen is reported to reproduce several stylized facts of financial markets, including realistic price distributions, heavy-tailed returns, and bursty event timing (Li et al., 4 Aug 2025).

1. Concept and motivation

ByteGen is designed for generative modeling of limit order book dynamics under the premise that market microstructure data are not naturally token-like objects. The paper argues that timestamps, prices, sizes, and order identifiers are high-precision numerical quantities for which tokenization introduces precision loss, artificial boundaries, destroyed ordinal structure, dataset-specific heuristics, and weak portability across assets and exchanges. In this formulation, byte space is treated as the native modeling domain for Level-3 market data rather than an inconvenient intermediate representation (Li et al., 4 Aug 2025).

The problem setting is defined by several difficulties: extremely long event sequences, nanosecond timestamps, non-uniform clustered arrivals, intricate dependencies across events and book state, and the need to preserve exact numeric precision. ByteGen is positioned against three major classes of prior approaches. First, stochastic, queueing, and agent-based models are described as analytically convenient but constrained by assumptions such as Markovian dynamics, independent Poisson arrivals, constant order sizes, or simplified queue behavior. Second, feature-engineered machine learning approaches are described as discarding fine-grained information through summaries such as spread, top-level depth, imbalance, and moving averages. Third, tokenized deep sequence models are described as inheriting both quadratic attention cost on ultra-long streams and tokenization dependence.

The conceptual claim of ByteGen is therefore not merely that it is autoregressive, but that it eliminates tokenization and feature engineering simultaneously. The model is intended to learn directly from exchange-like binary event encodings and to discover latent structure through the architecture rather than through handcrafted message segmentation.

2. Packed event representation in byte space

A central contribution is a fixed-length 32-byte packed binary format for each event. Every message is represented as

Event=Pack(ev_packed,exch_ts,price,quantity)\text{Event} = \text{Pack}(\text{ev\_packed}, \text{exch\_ts}, \text{price}, \text{quantity})

with four 8-byte fields: ev_packed as uint64, exch_ts as int64, price as float64, and quantity as float64 (Li et al., 4 Aug 2025).

Bytes Field Type
0–7 ev_packed uint64
8–15 exch_ts int64
16–23 price float64
24–31 quantity float64

The source fields selected from the Databento MBO schema are ts_event, action, side, price, size, and order_id. These are transformed into the packed format through bitwise composition. The first 64-bit word is defined as

ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}

so that the upper 32 bits hold order_id and the lower 32 bits hold ev. Within ev, bits 0–7 encode event type, with ADD_ORDER = 10, CANCEL_ORDER = 11, MODIFY_ORDER = 12, and FILL_EVENT = 13. Flag bits are assigned as EXCH_EVENT = 0x80000000, LOCAL_EVENT = 0x40000000, BUY_EVENT = 0x20000000, and SELL_EVENT = 0x10000000. Side is thus represented through the buy/sell flag bits rather than through a separate field.

The preprocessing logic described in the paper begins from raw MBO events, filters invalid records with zero price or zero quantity, constructs ev, packs ev_packed, serializes the four-field tuple into 32 bytes, concatenates events into a byte stream, and stores compressed arrays with NumPy compression. Reported compression ratios are typically 3–4x.

The paper repeatedly describes the representation as preserving all essential information and operating without loss of precision. A stricter reading, however, is narrower: losslessness is supported with respect to the chosen core event semantics rather than as a verbatim copy of the full original 64-byte native feed. The text itself makes that dependence visible through the selected field subset and through the assumption that order_id is representable in the upper 32 bits of ev_packed. The paper does not explicitly specify byte endianness.

3. Autoregressive formulation and H-Net architecture

ByteGen models a contiguous byte sequence rather than a sequence of symbolic market events. Each event occupies 32 bytes, and training samples are variable-length streams between 3,200 bytes and 10,240 bytes, corresponding to 100 to 320 events. Sequence boundaries are aligned to event boundaries, so samples begin and end on multiples of 32 bytes, but within each sample the model predicts one byte at a time and is not given explicit tokenized message boundaries (Li et al., 4 Aug 2025).

The sequence likelihood is written in standard autoregressive form:

p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T} p(x_t \mid x_{<t})

or equivalently over bytes btb_t,

p(b1:T)=∏t=1Tp(bt∣b<t).p(b_{1:T}) = \prod_{t=1}^{T} p(b_t \mid b_{<t}).

The main objective is next-byte cross-entropy together with a dynamic chunking ratio regularizer:

L=−1T∑t=1Tlog⁡p(bt∣b<t)+λ∑s=0S−1Lratio(s)\mathcal{L} = -\frac{1}{T}\sum_{t=1}^{T} \log p(b_t \mid b_{<t}) + \lambda \sum_{s=0}^{S-1} \mathcal{L}_{\text{ratio}^{(s)}}

with λ=0.01\lambda = 0.01. Inputs and labels are formed by shifting the byte stream, using s[:−1]s[:-1] as input and s[1:]s[1:] as target, and the implementation uses nested jagged tensors for variable-length batching.

The architecture is H-Net, described as a hybrid Mamba–Transformer hierarchy with dynamic chunking. It has three conceptual functions: a scanner for efficient processing of long raw byte streams, a segmenter that discovers variable-length chunks, and a reasoner that models long-range interactions over the compressed chunk sequence. Dynamic chunking is the mechanism by which ByteGen attempts to recover latent structure directly from bytes. Boundary probabilities are computed from adjacent hidden states through cosine similarity:

qt=Wqx^t,kt=Wkx^t,pt=12(1−cos_sim(qt,kt−1)),q_t = W_q \hat{x}_t, \quad k_t = W_k \hat{x}_t, \quad p_t = \frac{1}{2}\left(1 - \text{cos\_sim}(q_t, k_{t-1})\right),

with a boundary selected when ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}0. To preserve differentiability through boundary decisions, the model uses an EMA-style smoother,

ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}1

The stage compression ratios are reported as ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}2, and the empirical compression learned by dynamic chunking is reported to match the target, with Stage 0→1 at 4x and Stage 1→2 at 4x.

Within this hierarchy, Mamba-2 serves as the long-sequence scanner, while Transformer blocks serve as the higher-level reasoner once chunking has reduced the sequence length. The paper presents the Mamba component through continuous and discretized SSM equations and emphasizes input-dependent parameters ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}3, ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}4, and ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}5. Transformer blocks are introduced through standard self-attention equations. Stage blocks use a pre-norm residual form written generically as

ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}6

Three model scales are reported: a small model with 8M parameters, a base model with 124M parameters, and a large model with 1.5B parameters. Reported dropout is 0.0, and rotary dimension is specified, implying rotary positional encoding in the Transformer components.

4. Training corpus, optimization, and generation procedure

The training corpus consists of CME Bitcoin futures Level-3 Market-By-Order data from Databento, specifically the BTCX4 November 2024 contract over November 11–15, 2024. The dataset contains 34.2 million orderbook messages and occupies 1.02 GB in the 32-byte packed format. Reported market activity statistics include an average event rate of 79 events/sec over 24h, active periods of 100–400 events/sec, peaks above 1000 events/sec, a price range of $\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}$793,000, mean price of $87,506, and nanosecond timestamps. Inter-arrival times have median 0.271 ms and mean 13.311 ms; the paper uses the disparity between mean and median as evidence of burstiness and heavy tails (Li et al., 4 Aug 2025).

From a sampled 500,000-event table, event-type composition is reported as 36.8% MODIFY_ORDER, 31.5% ADD_ORDER, 31.3% CANCEL_ORDER, 0.2% FILL_EVENT, and 0.2% OTHER. This imbalance is later connected to generation bias.

Optimization details are specified as AdamW with learning rate ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}8, 1,000 warmup steps, cosine scheduling, batch size 16, weight decay 0.1, betas ev_packed=(order_id≪32)∣ev\text{ev\_packed} = (\text{order\_id} \ll 32) \mid \text{ev}9, epsilon p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T} p(x_t \mid x_{<t})0, gradient clipping 1.0, dropout 0.0, mixed precision in bfloat16, and distributed training through DDP and FSDP. The paper contains an internal inconsistency in training duration: one section reports 10,000 steps, another reports 20,000 steps, and convergence is described as occurring within 0.5 hour to 1 hour on 4 H100 GPUs. The safest factual summary is therefore that training used about 10k–20k steps on a 4-H100 cluster.

Generation is autoregressive: beginning from an initial byte context, the model samples the next byte repeatedly until the desired length is reached. The paper does not report temperature, top-p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T} p(x_t \mid x_{<t})1, top-p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T} p(x_t \mid x_{<t})2, or other decoding hyperparameters. Generated streams are parsed back into 32-byte event records and then unpacked into ev_packed, exch_ts, price, and quantity. One explicit validity constraint is enforced during generation:

p(x1:T)=∏t=1Tp(xt∣x<t)p(x_{1:T}) = \prod_{t=1}^{T} p(x_t \mid x_{<t})3

If a generated event violates temporal monotonicity, the system either regenerates with limited retries or applies minimal timestamp correction. Other structural validity issues, such as malformed event types, NaN floating-point payloads, or inconsistent flags, are not detailed in the paper.

5. Empirical behavior and evaluation

Evaluation is organized around three broad classes of diagnostics: price dynamics, market microstructure, and order flow characteristics. The reported metrics include price distribution similarity, price KL divergence, KS tests on event or return distributions, volatility clustering, autocorrelation patterns, inter-event time distributions, event type frequencies, order size distributions, bid-ask spread dynamics, order flow imbalance, order lifetime distributions, and fill rates. The paper does not provide a direct quantitative baseline table against specific competing models in the supplied text, so the evaluation is primarily real-versus-generated comparison rather than head-to-head benchmarking (Li et al., 4 Aug 2025).

The principal empirical claim is that ByteGen generates market-like event streams along several important dimensions. The paper reports price distributions as well-aligned, return distributions with matching heavy tails, and a Q-Q plot supportive of distributional similarity. Price KL divergence is reported as 0.023. Generated volatility is 12.6 bps versus 16.9 bps in real data, which indicates underestimation of volatility. Generated event rate is 154.3 events/sec versus 142.7 events/sec in real data. In a matched comparison window, generated data comprise 10,000 events over 64.796 seconds, compared with 9,329 real events over 64.748 seconds. Generated average spread is 2.8 bps versus 3.1 bps in real data.

Several limitations are explicit in the reported results. Generated order lifetime is 8.4 sec versus 11.2 sec in real data, and generated fill rate is 3.2% versus 8.7%. The paper also notes systematic event-type bias, with generated cancel orders at 47% versus 31% in real data, and states that the model produces fewer trades than observed historically. These figures identify execution realism and rare-event calibration as the clearest shortcomings.

The return-distribution summary further sharpens the picture. Mean return is 0.0001% in both generated and real data. Return standard deviation is 0.1568% generated versus 0.1590% real. Skewness is 0.3034 generated versus 0.5719 real. Kurtosis is 25.8557 generated versus 19.3124 real. This suggests that heavy tails are captured and somewhat exaggerated in kurtosis. Price statistics also show imperfect tail calibration: mean price is 91,185.46 generated versus 91,052.51 real, price standard deviation is 115.25 versus 154.43, minimum price is 89,276 versus 90,130, and maximum price is 92,785 versus 92,055. Event timing is described as qualitatively realistic, with inter-event times following a power-law-like distribution, though slight tail differences remain. The main table reports an Event KS Statistic of 0.187.

6. Limitations, interpretation, and broader research context

ByteGen’s strengths are stated in terms of representation and modeling scope: it avoids tokenization bias, preserves exact numerical precision by modeling raw float64 and int64 bytes, eliminates handcrafted microstructure features, learns latent segments through dynamic chunking, and uses Mamba components to mitigate long-sequence scaling constraints. The practical uses proposed for a sufficiently calibrated version include market simulation, strategy backtesting, stress testing, reinforcement-learning environments, and market impact studies (Li et al., 4 Aug 2025).

The limitations are equally central to its interpretation. Byte-level autoregression is computationally expensive because each event contributes 32 prediction steps. The fixed 32-byte schema is efficient but format-specific. Rare but important events, especially fills and trades, remain difficult. Temporal monotonicity is enforced during generation, which means not all structural market constraints are learned internally. Evaluation is weakened by the absence of extensive head-to-head baselines and formal ablations in the supplied text. Some implementation details that matter for exact reproducibility remain unspecified, including byte endianness, the precise byte embedding design, the exact stage/block allocation for each model size, the train/validation/test split, and the sampling hyperparameters.

Several common misconceptions can therefore be addressed directly. “Tokenizer-free” does not mean “constraint-free,” because timestamp monotonicity is enforced post hoc during generation. “Lossless” does not mean a verbatim copy of the original native feed, but preservation of the selected core event semantics under the stated packing assumptions. “Competitive performance” should also be read narrowly, because the paper does not provide strong explicit baseline methodology in the provided text.

A broader methodological connection can be drawn to earlier work on direct low-level representations. One such line studies automatic generation of Forth bytecode from input-output tests, using linear instruction sequences, probabilistic generation, Markov generation, and a genome-like mechanism that promotes frequent subprograms into reusable words (Khashin et al., 2018). That work is not about limit order books, but it is relevant as a conceptual precursor in its commitment to machine-near representations and learned structural reuse. This suggests a broader research tendency, across otherwise distinct domains, toward eliminating higher-level symbolic preprocessing in favor of byte- or instruction-level modeling.

In that larger context, ByteGen marks a specific shift in financial sequence modeling: it treats the market feed as native binary data and delegates segmentation to a learned hierarchical architecture. The reported results indicate that this is sufficient to recover several high-level stylized facts of market behavior, while the remaining deficiencies concentrate in rare-event modeling, execution realism, and evaluation rigor.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to ByteGen.